
Share
A routine employment claim lands on a Monday, with one custodian and one preservation notice attached. It should take an afternoon to collect and a week to move past. Then the volume report arrives, and the afternoon turns into most of a quarter.
Nine years of mail, a shared drive nobody has pruned since the last migration, and a Slack workspace with no retention rule at all. Nothing actually failed inside that matter, and every step of it was handled correctly. What failed happened in 2019, when the company quietly chose not to decide how long any of that data should live.
That non-decision is now the largest single line item in the case budget. It is also the part of information governance that most definitions manage to skip entirely. Every organization already has a program of some kind, and yours is simply whatever your systems do when nobody sets a rule.
The unwritten default in almost every environment is to keep everything, forever, by accident. So the useful question is not whether your organization has an information governance program. It is whether that program can delete anything at all and defend the decision afterward.
Information governance is the coordinated set of rules that decides how an organization creates, stores, protects, retains, and destroys information. It spans records management, data privacy, information security, and eDiscovery. The purpose is to meet legal obligations and limit risk while keeping information genuinely useful to the business.
The Sedona Conference frames it as a coordinated, interdisciplinary approach to compliance requirements, information risk, and information value. That framing matters because it places four separate disciplines under a single roof. Most organizations, however, still run those four disciplines in four separate rooms.
Legal owns preservation, IT owns storage, security owns access, and privacy owns consent and collection limits. Nobody owns the end of the lifecycle, which happens to be where the expensive decisions actually sit. That ownership gap is why governance programs produce documents far more reliably than they produce outcomes.
It also helps to separate three terms that get used interchangeably and mean different things. An information governance policy is the document, while an information governance strategy is the plan sitting behind it. The program is what actually happens to a file on its last useful day.

Records management is a discipline inside information governance rather than a synonym for it. It classifies records and applies a records retention policy to each defined class. Information governance sits above that work and answers the harder questions about privacy duties, security controls, and discovery exposure.
Data governance is a third and separate thing again, and the terms are often confused. It concerns whether your data is accurate, consistent, and usable as a business asset. Information governance concerns whether you are legally permitted to keep that data at all.
Information governance matters because every byte you retain is a byte you may have to collect, review, produce, secure, and explain. Retention feels cheap right up until litigation, a breach, or a regulator makes it expensive. Governance is how an organization prices that risk in advance instead of discovering it under deadline.
Without it, the cost surfaces on somebody else’s timeline, and that somebody is usually the legal team. Storage decisions made quietly by default in IT become discovery problems owned outright by counsel. The bill arrives years later, in a different department, with no realistic way to appeal it.
Information governance occupies the leftmost stage of the EDRM model, before anything is identified or preserved. That placement is easy to misread as a first step you complete and then move past. It is closer to the foundation that every later stage quietly stands on.
The framework view of how stakeholders divide this work is covered in our guide to the IGRM model. This piece takes a different route and looks at what a program actually produces. Both views matter, but only one of them ever shows up in a sanctions motion.
Ask a room of legal ops leaders when their organization last deleted a category of data on purpose and documented the reasoning. The pause that follows is the most honest answer available anywhere in this field. Almost every program can produce a policy on request, and almost none can produce a deletion.
Classification schemes, access controls, and a records retention schedule are all inputs to governance. Defensible disposition is the output, and it is the only part a court or a regulator ever inspects. Everything upstream of that is preparation nobody outside the organization will see.
The Sedona Conference gives disposition its own principle for precisely this reason. Its commentary observes that organizations still struggle to make and execute effective disposition decisions, years after the problem was first named. The difficulty is operational rather than conceptual, which is why policy alone never resolves it.
Data disposition is also the hardest part of governance to do halfway. Deleting selectively requires knowing what you hold, what obligation attaches to it, and whether a hold is currently active. Most programs stall at the first of those three questions and never reach the third.
So a better working test replaces the usual one, and it is uncomfortably simple to apply. Instead of asking whether you have an information governance policy, ask what your program destroyed in the last twelve months. If the answer is nothing, the policy is describing a company you do not actually have.
Storage is the cheapest part of keeping data, which is precisely why over retention hides so effectively. The real invoice arrives much later and lands in an entirely different budget line. Sedona puts it plainly, noting that failing to dispose of information that no longer adds value increases both the cost and the risk of meeting discovery obligations.
Dark data is the usual culprit in most environments. It is information an organization holds but cannot see, describe, or justify, and it still gets collected the moment a custodian falls in scope. ROT data behaves the same way, since redundant, obsolete, and trivial content adds nothing to the business and everything to the gigabyte count.

That gigabyte count is what sets your review budget, and it compounds at the worst possible moment. Culling that could have happened across three quiet years now has to happen under a discovery deadline, which is where early case assessment absorbs the difference. Governance decisions deferred slowly upstream become emergencies handled quickly downstream.
Privacy exposure follows exactly the same curve. Data retained without purpose is data that can be breached, and data you must locate when a data subject access request arrives with a statutory clock attached. The same over retention that inflates review volume also widens every privacy obligation the organization carries.
Most records retention schedules were written for a company that no longer exists in that form. They assume email, file shares, and paper, with a records manager sitting between the business and the archive. Three failure patterns show up repeatedly across organizations, and none of them are drafting problems.
A retention schedule that lives in a policy binder governs precisely nothing. Unless it is configured inside Microsoft 365, Google Workspace, Slack, and every SaaS tool in active use, the platform default quietly wins. Most teams discover that mismatch during a collection rather than during an audit.
Unstructured data governance is where retention schedules usually collapse under their own weight. Chat threads, meeting recordings, and shared drive folders rarely map cleanly to any defined record class. Anything that resists classification defaults to indefinite retention, which is a decision nobody consciously made.
Every new platform enters the environment without a retention rule attached to it. Adoption moves in weeks while governance review moves in quarters, and that gap tends to become permanent. By the time anyone revisits the question, the tool already holds two years of unclassified content.

The result is a documented data retention policy describing intent, alongside a live environment doing something else entirely. Discovery is usually where those two versions are formally introduced to each other. That introduction rarely goes well for the party who wrote the policy.
This is the tension most information governance content avoids, and it is now the defining problem in the field. Twenty states have comprehensive consumer privacy laws in effect in 2026, with no federal law overriding them. Data minimization sits close to the center of nearly all of those statutes.
Maryland has pushed that standard the furthest of any state so far. Its Online Data Privacy Act took effect on October 1, 2025 and limits collection to what is reasonably necessary and proportionate to a product the consumer actually requested. Privacy law, in other words, is pushing organizations toward collecting less and deleting sooner.
Federal Rule of Civil Procedure 37(e) pushes firmly in the opposite direction. It applies when electronically stored information that should have been preserved is lost because reasonable steps were not taken. Where a party acted with intent to deprive another of that evidence, a court may instruct the jury to presume it was unfavorable, or impose spoliation sanctions as severe as default judgment.
The two duties genuinely conflict, and the Sedona Conference addresses that conflict head on. Principle 7 asks organizations to reconcile competing obligations in good faith, and Principle 8 asks reviewing courts to weigh that effort fairly. Good faith here means a documented reasoning trail, not a confident instinct.
In practice, a single sentence resolves most of the conflict, and most retention policies do not contain it. A legal hold suspends the retention schedule for the data it covers, and nothing else does. That rule is why the boundary between legal hold, preservation, and collection has to be written down rather than assumed.
Ungoverned data is rarely the data an organization already knows about. It is whatever entered the environment faster than the policy could follow, and generative AI has widened that gap considerably. Five sources account for most of the unmanaged exposure in a typical Microsoft environment.
Microsoft 365 Copilot prompts and responses are stored as individual items inside the user’s Exchange mailbox. They appear in broad mailbox searches unless somebody explicitly filters them out. That means they are already sitting in your collections, whether or not anyone planned for them to be.
Purview retention policies and eDiscovery holds do not currently apply to Copilot memory. Deleting a prompt and its response does not remove the memory derived from that conversation. Memory therefore has to be identified and handled as a separate step, which most retention programs have never accounted for.
Copilot Pages are not stored inside Exchange mailboxes at all. A hold placed on a custodian mailbox does not reach them, and the SharePoint Embedded site holding them has to be added to the hold policy by URL. A mailbox hold that looks complete can therefore miss an entire artifact class.
Transcript and summary tools now attach to calls across most organizations, often without procurement review. They create a durable written record of conversations that previously left no artifact behind at all. Every one of those records is discoverable, and almost none of them appear on a retention schedule.
Slack and Teams remain the least governed high value sources in most environments. Retention in both platforms is usually a default that nobody selected deliberately. The decisions that eventually end up in litigation increasingly happen in exactly these channels.

None of these sources are exotic, and that is precisely what makes them dangerous. They are ordinary tools that arrived without a retention decision attached, which is how governance gaps accumulate without anyone approving them. Microsoft sets out where each Copilot artifact lives and what a hold does and does not reach in its own Purview eDiscovery guidance.
Before rebuilding anything, run these five questions against your current state. Each one has a clear failure mode, and each fix is considerably smaller than a full program overhaul. Answering them honestly takes an afternoon and reveals more than a formal maturity assessment usually does.
If the answer is nothing, you have a storage habit rather than a governance program. Start by disposing of one low risk, high volume category and documenting the reasoning behind it. That single act creates the precedent everything else can follow.
Compare your records retention schedule against the live configuration in your three largest repositories. Any gap between the two is what a court will eventually see, not the document itself. Close that gap or formally accept it, but never leave it undocumented.
Defensible disposition needs a named approver and a written trail behind every decision. If nobody owns that decision, it will keep being deferred by default until a matter forces the issue. Ownership is usually the cheapest fix on this entire list.
Test this rather than assume it, because assumptions are where holds quietly fail. Place a hold on a sample custodian and confirm that deletion genuinely stops across mail, chat, and cloud storage. Teams are frequently surprised by which system keeps deleting anyway.
List every tool added in the last three years and mark those with no retention setting configured. That list is your governance backlog, already usefully ranked by risk. Work down it in order of data volume rather than in order of discovery.

A first pass at information governance should produce artifacts rather than activity. Four outputs are enough to change the trajectory of an entire program. Each one is concrete enough that an executive sponsor can verify it was genuinely completed.
That last item carries more weight than the other three combined. An information governance strategy earns credibility the first time it disposes of something and can explain exactly why. Everything before that point is preparation, and preparation has never been a defense.
The employment claim from the opening was never really about nine years of mail. It was about a decision nobody made in 2019, surfacing on a deadline six years later. That is how information governance always presents itself, invisible while it is cheap and unavoidable once it turns expensive.
The organizations that handle discovery well are rarely the ones with the most sophisticated policy documents. They are the ones that decided early, and in writing, what they would keep and what they would let go. The document itself matters far less than the record of decisions sitting behind it.
Venio brings legal hold, early case assessment, review, and production onto a single platform. The governance decisions you make upstream then still hold when a matter actually begins. Talk to our eDiscovery Expert to see what that looks like across your own data sources.
A data retention policy is the written rule setting how long each category of information is kept and when it is destroyed. It usually pairs with a records retention schedule that assigns specific periods by record type. It only works once it is configured in the systems that actually hold the data.
Retention periods depend on record type, industry, and the jurisdictions you operate in, so no single number applies. Tax, employment, and regulated financial records carry statutory minimums your schedule has to reflect. Everything beyond those minimums is a business decision you should make deliberately rather than by default.
Responsibility is shared across legal, IT, security, privacy, and records management, which is why so many programs stall without a named owner. Most mature programs appoint an executive sponsor alongside a cross functional steering group. Sedona also recommends the program stay independent enough that no single department drives every decision.
Yes, and the gaps there are considerably wider than most programs assume. Copilot prompts and responses sit in the user mailbox and appear in eDiscovery collections, while Copilot memory falls outside standard retention and hold coverage. Any tool that generates or stores content belongs inside your governance scope from day one.