Metadata is data that describes other data. For an electronic file, metadata records details such as who created it, when it was last modified, and where it was stored, rather than what the file says. In eDiscovery, metadata is extracted during processing and used to authenticate evidence and reconstruct timelines.
Share
Metadata is descriptive information carried alongside a file rather than inside its visible content. The standard metadata definition, data about data, is accurate but incomplete. What makes metadata useful is that it answers questions the document itself cannot: a memo shows what was written, while its metadata shows who wrote it, when, on what system, and whether it was altered afterward.
Courts have described metadata as embedded information recording the history, tracking, and management of an electronic document.
Metadata attaches to every form of electronically stored information, including email, office documents, chat messages, and files pulled from cloud platforms. Some of it is visible to the user, such as a tracked change or a comment. Most of it is not, which is what makes metadata both evidence and a hazard.
Courts recognize three types of metadata. In Aguilar v. Immigration and Customs Enforcement, a federal court set out the categories that eDiscovery practitioners now apply.
Processing platforms generate a fourth layer, sometimes called derived metadata. Hash values, document IDs, and parent-attachment relationships do not originate with the file, but they travel with the record through review and production.
Library and information science uses a different set of categories, descriptive, structural, and administrative. Those terms describe the same underlying concept but are not the framework courts apply in discovery disputes.
Metadata collection begins at the source. When files are gathered from a custodian's mailbox, laptop, or cloud account, the collection method determines how much survives. Copying files by hand through an operating system overwrites access dates and can reset authorship, which is why defensible collection uses forensic tools that capture the file and its metadata together.
Metadata extraction happens next, during the processing stage of the EDRM. Processing software opens each file, reads its metadata into structured fields, and indexes those fields alongside the extracted text. Those fields are what make deduplication, email threading, and targeted search possible across a full dataset.
A standard production field set includes:
Hash values serve a distinct role. Any change to a file produces a different hash, which makes it the standard mechanism for verifying that evidence has not been altered.
Preservation runs alongside all of this. Once a duty to preserve is attached, it reaches the metadata as well as the file, because a document with reset dates or stripped authorship carries far less evidentiary weight than the same document collected intact.
Metadata is discoverable when it is relevant to a claim or defense and not privileged. Its practical value falls into three areas: authenticating a document, establishing when events occurred, and making large volumes of data searchable and sortable.
Form of production determines whether metadata reaches the other side. Under Federal Rule of Civil Procedure 34(b)(1)(C), a requesting party may specify the form. Absent that, Rule 34(b)(2)(E)(ii) requires production in the form the information is ordinarily maintained or in a reasonably usable form. In Blevins-Clark v. Beacon Communities, a court held that native files produced as PDFs stripped of metadata did not meet that standard and ordered re-production.
Image productions handle this differently. When files are converted to TIFF or PDF, the metadata travels separately in a load file, a delimited text file mapping each field to its corresponding document. The image itself carries no native metadata, which is also why OCR is needed to make those pages searchable.
Timing is the trap practitioners hit most often. Courts have ordered metadata produced where it was specifically requested in the initial document request and the producing party had not yet produced the documents in any form. A request made after production has already occurred in another form is considerably weaker.
Venio eDiscovery extracts and indexes metadata fields during processing, keeping custodian, date, and hash values consistent from collection through production. Book a demo to see it on your own data.
Metadata is data that describes other data. For a file, it records details such as who created it, when it was last modified, and where it was stored, rather than what the file contains.
System metadata, substantive or application metadata, and embedded metadata. System metadata is generated by the operating system, substantive metadata by the application used to create the file, and embedded metadata is entered into a file but not displayed on screen.
Yes, when it is relevant to a claim or defense and not privileged. Courts have generally ordered metadata produced where it was specifically requested in the initial document request and the documents had not already been produced in another form.
Copying files through an operating system, converting them out of native format, or producing them as images without a load file can overwrite or strip metadata. Forensic collection tools capture the file and its metadata together to prevent this.
Metadata is data that describes other data. Learn the metadata definition, the types of metadata courts recognize, and how metadata extraction works i
Read MoreEarly case assessment (ECA) surfaces a matter's facts, risk, and cost early. Learn how ECA differs from review and guides the settle or litigate call.
Read MoreA legal hold, or litigation hold, preserves relevant data once litigation is anticipated. Learn how the process and the hold notice work in eDiscovery
Read MoreThe duty to preserve is the legal obligation to protect evidence once litigation is reasonably anticipated. Learn how preservation of evidence works i
Read More