Data processing is the eDiscovery stage that converts a raw, collected ESI set into an indexed, searchable, review-ready dataset. It sits between collection and review in the EDRM, and it covers ingestion, exception handling, culling, metadata and text extraction, OCR, and format conversion, the six steps that determine how much data, and how much cost, ultimately reaches attorney review.
Share
A collected dataset is not a reviewable one. What comes out of collection is a mix of native files, mail containers, forensic disk images, and cloud exports. Each one is still in whatever structure it was created in. Data processing is the stage that turns that raw mix into something a review platform can index and a person can actually work through.
Processing is the fourth stage of the EDRM, it's also the highest-leverage stage in the model. The decisions made here, how well exceptions get resolved, how much gets culled, whether metadata is captured correctly, set the volume that reaches review. And volume is what drives cost. Good eDiscovery data processing pays for itself by the time a matter reaches review. Sloppy processing does the opposite, and the damage usually doesn't surface until review is already underway.
Processing runs as a sequence, not a single action. Six steps carry a dataset from raw collection to review-ready:
Deduplication, OCR, and native versus image formatting each have enough depth to warrant their own explanation elsewhere in this glossary. This page is about how all six steps work together as one pipeline.
Volume reduction numbers get most of the attention in processing conversations, but the number that actually protects a case is how thoroughly the job was checked before it moved downstream. A defensible processing workflow runs three layers of quality control, not one.
Item counts going into processing should match items coming out, accounting for every exception. If the numbers don't reconcile, something was dropped silently, and nobody should find that out during review.
A random sample of extracted text and metadata gets spot-checked against the source files, catching systemic extraction problems, like a metadata field that mapped incorrectly across an entire custodian's mailbox, before they propagate through the whole dataset.
The exception log, every corrupted, encrypted, or password-protected file the engine couldn't fully process, goes to the attorney or project manager overseeing the matter for an explicit decision. Pursue a password, seek a replacement file, or document why the file is being excluded. An exception report that nobody reads defeats its own purpose.
Skipping any of these three doesn't necessarily mean anything goes wrong. It means nobody would know if it did, until it surfaces in a discovery dispute instead of a QC report.
Under Federal Rule of Civil Procedure 26(g)(1)(B), an attorney's signature on a discovery response certifies that, after reasonable inquiry, the response is complete and consistent with the rules, a duty sometimes called the stop and think rule. An unresolved exception report or an unsupervised processing job doesn't automatically break that certification, but it removes the evidence that reasonable inquiry actually happened.
In Grullon v. Lewis, 2025 WL 1693425 (S.D.N.Y. June 17, 2025), a court found a Rule 26(g) violation where client-run searches proceeded with minimal oversight from counsel, a reminder that supervision of how data gets processed and searched, not just the final production, is part of what the certification actually covers.
Venio runs ingestion, exception handling, culling, and OCR in a single processing engine, with reconciliation and exception reporting built into every job rather than bolted on afterward. Book a demo to see it process your own dataset.
Data processing is the eDiscovery stage that converts a raw, collected ESI set into a searchable, indexed, review-ready dataset. It includes ingestion, exception handling, culling, metadata and text extraction, OCR, and format conversion, and it sits between collection and review in the EDRM.
ESI processing is another name for eDiscovery data processing, the stage where electronically stored information (ESI) is culled, converted, and indexed so it can be searched and reviewed.
They're logged as exceptions rather than silently skipped. A defensible processing workflow tracks every exception, corrupted files, encrypted containers, password-protected documents, so the legal team can decide whether to pursue a password, a replacement file, or another remediation path, instead of losing that data without knowing it.
Because a volume-reduction number says nothing about whether the reduction was accurate. A processing job with impressive culling stats but no reconciliation, sampling, or exceptions review could still be silently dropping responsive data, the QC layer is what actually makes the reduction defensible.
Data processing converts raw ESI into a reviewable, indexed dataset. See the six-step pipeline, the QC checks that make it defensible, and why it driv
Read MoreCulling reduces an ESI dataset to what's relevant and proportional before review. Learn how eDiscovery culling works, its techniques, and why it must
Read MoreNative format keeps a file's original structure and metadata. Image format (TIFF/PDF) is a flattened picture of it. Learn when eDiscovery uses each.
Read MoreMetadata is data that describes other data. Learn the metadata definition, the types of metadata courts recognize, and how metadata extraction works i
Read More