Home

    /

    Glossary

    /

    What Is Data Processing in eDiscovery?

    What Is Data Processing in eDiscovery?

    August 5, 2026
    Reading Time :
    Definition

    Data processing is the eDiscovery stage that converts a raw, collected ESI set into an indexed, searchable, review-ready dataset. It sits between collection and review in the EDRM, and it covers ingestion, exception handling, culling, metadata and text extraction, OCR, and format conversion, the six steps that determine how much data, and how much cost, ultimately reaches attorney review.

    On this page

    Share

    A collected dataset is not a reviewable one. What comes out of collection is a mix of native files, mail containers, forensic disk images, and cloud exports. Each one is still in whatever structure it was created in. Data processing is the stage that turns that raw mix into something a review platform can index and a person can actually work through.

    Processing is the fourth stage of the EDRM, it's also the highest-leverage stage in the model. The decisions made here, how well exceptions get resolved, how much gets culled, whether metadata is captured correctly, set the volume that reaches review. And volume is what drives cost. Good eDiscovery data processing pays for itself by the time a matter reaches review. Sloppy processing does the opposite, and the damage usually doesn't surface until review is already underway.

    The Data Processing Pipeline, Step by Step

    Processing runs as a sequence, not a single action. Six steps carry a dataset from raw collection to review-ready:

    1. Ingestion And Format Identification: The raw collection is loaded, and every file gets identified by type. Native documents get sorted from mail containers, and mail containers from system files. You can't filter what you haven't correctly identified yet, which is why this happens first. 
    2. Exception Handling: Files the engine can't fully open, corrupted documents, password-protected files, encrypted containers, get logged in an exceptions report instead of silently dropped. This is the step most processing explainers skip past. It's also the one with the highest downside. An unresolved corrupted mail container can mean losing tens of thousands of emails without anyone noticing, because a file that never processes never shows up in search.
    3. Culling: Date range, file type, custodian, and domain filters narrow the dataset. Deduplication and DeNISTing run here too. DeNISTing strips known system files using the NIST National Software Reference Library (NSRL). This is usually where the biggest volume drop happens. 
    4. Metadata And Text Extraction: Custodian, dates, author, email headers, and each document's underlying text get pulled and indexed. This is what makes the dataset searchable. 
    5. OCR: Scanned pages and photographed documents get a text layer, so they join the searchable set instead of sitting invisible to keyword search.
    6. Format Conversion And Load File Creation: The dataset is prepared for review and production, native, TIFF, PDF, or a near-native hybrid, along with the load files that carry Bates numbers and metadata for whichever format the matter calls for. 

    Deduplication, OCR, and native versus image formatting each have enough depth to warrant their own explanation elsewhere in this glossary. This page is about how all six steps work together as one pipeline.

    The Three-Layer QC Check Every Processing Job Needs

    Volume reduction numbers get most of the attention in processing conversations, but the number that actually protects a case is how thoroughly the job was checked before it moved downstream. A defensible processing workflow runs three layers of quality control, not one.

    1. Automated Reconciliation

    Item counts going into processing should match items coming out, accounting for every exception. If the numbers don't reconcile, something was dropped silently, and nobody should find that out during review.

    2. Statistical Sampling

    A random sample of extracted text and metadata gets spot-checked against the source files, catching systemic extraction problems, like a metadata field that mapped incorrectly across an entire custodian's mailbox, before they propagate through the whole dataset.

    3. Sign-Off On The Exceptions Report 

    The exception log, every corrupted, encrypted, or password-protected file the engine couldn't fully process, goes to the attorney or project manager overseeing the matter for an explicit decision. Pursue a password, seek a replacement file, or document why the file is being excluded. An exception report that nobody reads defeats its own purpose.

    Skipping any of these three doesn't necessarily mean anything goes wrong. It means nobody would know if it did, until it surfaces in a discovery dispute instead of a QC report.

    Why Processing Quality Has Legal Teeth

    Under Federal Rule of Civil Procedure 26(g)(1)(B), an attorney's signature on a discovery response certifies that, after reasonable inquiry, the response is complete and consistent with the rules, a duty sometimes called the stop and think rule. An unresolved exception report or an unsupervised processing job doesn't automatically break that certification, but it removes the evidence that reasonable inquiry actually happened. 

    In Grullon v. Lewis, 2025 WL 1693425 (S.D.N.Y. June 17, 2025), a court found a Rule 26(g) violation where client-run searches proceeded with minimal oversight from counsel, a reminder that supervision of how data gets processed and searched, not just the final production, is part of what the certification actually covers.

    Key Takeaways

    • Data processing, also called ESI processing, is the EDRM stage that turns a raw collection into an indexed, reviewable dataset.
    • It runs as a six-step pipeline, ingestion and format identification, exception handling, culling, metadata and text extraction, OCR, and format conversion.
    • Exception handling is the step most likely to be underweighted, and the one with the highest downside. An unresolved corrupted container can mean losing a significant share of a custodian's data without anyone noticing.
    • A defensible processing job runs three layers of QC: automated reconciliation, statistical sampling, and explicit sign-off on the exceptions report, not just a volume-reduction number.
    • Under Rule 26(g)(1)(B), counsel's certification of a discovery response depends on reasonable inquiry, which extends to supervising how data was processed, not only what was ultimately produced.

    Venio runs ingestion, exception handling, culling, and OCR in a single processing engine, with reconciliation and exception reporting built into every job rather than bolted on afterward. Book a demo to see it process your own dataset.

    Frequently Asked Questions

    What is data processing in eDiscovery?

    Data processing is the eDiscovery stage that converts a raw, collected ESI set into a searchable, indexed, review-ready dataset. It includes ingestion, exception handling, culling, metadata and text extraction, OCR, and format conversion, and it sits between collection and review in the EDRM.

    What is ESI processing?

    ESI processing is another name for eDiscovery data processing, the stage where electronically stored information (ESI) is culled, converted, and indexed so it can be searched and reviewed.

    What happens to corrupted or password-protected files during processing?

    They're logged as exceptions rather than silently skipped. A defensible processing workflow tracks every exception, corrupted files, encrypted containers, password-protected documents, so the legal team can decide whether to pursue a password, a replacement file, or another remediation path, instead of losing that data without knowing it.

    Why does processing quality matter more than the reduction percentage?

    Because a volume-reduction number says nothing about whether the reduction was accurate. A processing job with impressive culling stats but no reconciliation, sampling, or exceptions review could still be silently dropping responsive data, the QC layer is what actually makes the reduction defensible.

    Explore More Glossary Terms

    What Is Data Processing in eDiscovery?

    Data processing converts raw ESI into a reviewable, indexed dataset. See the six-step pipeline, the QC checks that make it defensible, and why it driv

    Read More

    What Is Culling in eDiscovery?

    Culling reduces an ESI dataset to what's relevant and proportional before review. Learn how eDiscovery culling works, its techniques, and why it must

    Read More

    What Is Native vs. Image Format in eDiscovery?

    Native format keeps a file's original structure and metadata. Image format (TIFF/PDF) is a flattened picture of it. Learn when eDiscovery uses each.

    Read More

    What Is Metadata (eDiscovery)?

    Metadata is data that describes other data. Learn the metadata definition, the types of metadata courts recognize, and how metadata extraction works i

    Read More

    Protect your evidence before spoliation becomes a problem.

    See how Venio Legal Hold helps your team issue, track, and document defensible holds in minutes.

    No credit card required • Free product tour available