Home

    /

    Glossary

    /

    What Is Native vs. Image Format in eDiscovery?

    What Is Native vs. Image Format in eDiscovery?

    August 3, 2026
    Reading Time :
    Definition

    Native format is a file kept in the application that created it, such as a .docx, .xlsx, or .msg, with its original metadata and formatting intact. Image format is a flattened picture of that file, typically a TIFF or PDF, that looks the same on every screen but carries no native metadata of its own. eDiscovery productions use one, the other, or both, depending on what the case requires.

    On this page

    Share

    Every document produced in discovery exists in one of these two states before it reaches the other side. Native format means the file stays in the format it was created in, a Word document is still a .docx, an email is still a .msg or .eml, a spreadsheet is still an .xlsx with its formulas live underneath the displayed values.

    Image format means the file has been converted into a static picture of itself, almost always a TIFF (Tagged Image File Format) or a PDF. A TIFF or image-format PDF shows exactly what the page looked like, but it is a flattened image, not the underlying file. Open a native Excel file and the formulas are still there to inspect; open its TIFF and you see only the numbers a cell displayed at the moment of conversion.

    Neither format is inherently correct. Which one a matter uses, and for which document types, is a decision that belongs in the ESI protocol before production starts, not something to work out after documents are already moving. The choice sits inside the Production stage of the EDRM, but it gets set much earlier, at the meet-and-confer, based on what document review will actually need.

    How Native and Image Production Work

    Native production hands over the file as-is. Nothing is converted, so nothing is lost, embedded formulas, tracked changes, speaker notes, and the full metadata record travel with the file. The tradeoff is that native files are hard to redact reliably and hard to Bates-stamp in a way that is visible on the document itself, since the content sits inside proprietary application structures rather than on a fixed page.

    Image production converts each document to TIFF or PDF at the time of document review or production. This flattening is what makes redaction possible: a reviewer can black out a passage on the image, and because the image is a picture rather than editable text, the redacted content cannot be recovered the way it sometimes can be from a poorly redacted native file. Image production also makes Bates numbering straightforward, since a sequential number can be stamped directly onto each page.

    Because a TIFF or PDF carries no native metadata of its own, image productions travel with two companion files:

    • OPT file: Opticon cross-reference file is a delimited text file that maps each Bates-numbered image to its place in the document set, so a review platform can rebuild page order and document breaks.
    • DAT file: A Data file carries the metadata and extracted text fields, custodian, dates, author, subject line, that the image itself cannot hold.

    Image-only files also need OCR (optical character recognition) run against them before the text becomes searchable at all, since a TIFF is pixels, not characters, until that layer is added.

    Why Native vs. Image Format Matters

    The choice affects cost, defensibility, and how much of the document survives.

    Redaction is the most common reason a document moves from native to image. Privileged or sensitive content can be reliably redacted on a flattened image in a way that is difficult to guarantee on a native file, where hidden text, speaker notes, or tracked changes can survive a redaction that only addresses the visible surface.

    Cost runs the other direction. TIFF conversion, OCR, and load file generation add a per-page or per-GB processing cost that native production skips entirely. At volume, that difference is real money, and image files also consume more storage than the natives they were converted from.

    Rule compliance sits underneath both. Under Federal Rule of Civil Procedure 34(b)(1)(C), the requesting party may specify the form of production. If the responding party wants to produce something else, Rule 34(b)(2)(D) requires a specific objection that states the form it intends to use instead, a silent switch to TIFF when native was requested is not a valid response. 

    A related provision, Rule 34(b)(2)(E)(iii), also means a party generally cannot be made to produce the same ESI in both forms unless the parties agree to it, so getting the format decision right before production begins matters more than it might seem.

    Native vs. TIFF vs. Near-Native: Choosing a Format

    Most ESI protocols do not pick one format for an entire matter. They set defaults by document type and let exceptions get resolved by agreement.

    • Native: It is typical for spreadsheets, presentations, audio, video, and any file that loses meaningful content when flattened, a TIFF of an Excel file with hundreds of formula-driven tabs is close to useless for analysis.
    • TIFF/Image: It  is typical for emails and standard documents that need redaction, Bates stamping, or a consistent, non-editable appearance across every reviewer's screen.
    • Near-native (hybrid) Production: It combines the two. Documents go out as TIFF or PDF images, each carrying a load file with full metadata and extracted text, so the production looks like an image set but preserves most of what a pure native production would have carried. This is now the working default in many ESI protocols precisely because it balances redaction control against metadata loss.

    Key Takeaways

    • Native format keeps a file in its original application format, with metadata and structure intact; image format is a flattened TIFF or PDF picture of it.
    • Image production enables reliable redaction and Bates stamping; native production preserves metadata and content that doesn't survive flattening, like formulas and speaker notes.
    •  TIFF and PDF images carry no native metadata of their own, so they travel with OPT and DAT load files to carry Bates numbers, metadata, and extracted text.
    • Under FRCP 34(b)(1)(C), the requesting party may specify the form of production, and a responding party that wants to use a different form must object with specificity under Rule 34(b)(2)(D).
    • Most matters don't choose one format for everything; ESI protocols typically default to native for spreadsheets and complex files, and TIFF or near-native for standard documents and email.

    Venio Production supports native, TIFF, PDF, and near-native output from a single platform, with automated Bates numbering, load file generation, and AI-assisted validation built into the workflow. Book a demo to see it handle your own production formats. 

    Frequently Asked Questions

    What is the difference between native and image format in eDiscovery?

    Native format is a file left in the format it was created in, with its metadata and structure intact. Image format is a flattened picture of that file, usually a TIFF or PDF, that displays the same way on any screen but carries no metadata of its own.

    Can native files be redacted?

    Not reliably. Redacting a native file risks leaving hidden text, tracked changes, or speaker notes intact underneath the visible redaction. Most teams convert a document to image format before redacting it, since a flattened image cannot expose content that sits beneath the surface.

    What is an OPT file used for?

    An OPT file is the image load file that accompanies a TIFF or PDF production. It tells a review platform how the individual page images map back to complete documents, preserving page order and document breaks that the images alone don't carry.

    Does FRCP require production in native format?

    No single format is required by default. Under Rule 34(b)(1)(C), the requesting party may specify a form; if it doesn't, Rule 34(b)(2)(E)(ii) requires production in the form the ESI is ordinarily maintained or in a reasonably usable form. Courts have generally required a producing party that wants to use a different format to raise a specific, timely objection rather than substitute one unilaterally.

    Explore More Glossary Terms

    What Is Culling in eDiscovery?

    Culling reduces an ESI dataset to what's relevant and proportional before review. Learn how eDiscovery culling works, its techniques, and why it must

    Read More

    What Is Native vs. Image Format in eDiscovery?

    Native format keeps a file's original structure and metadata. Image format (TIFF/PDF) is a flattened picture of it. Learn when eDiscovery uses each.

    Read More

    What Is Metadata (eDiscovery)?

    Metadata is data that describes other data. Learn the metadata definition, the types of metadata courts recognize, and how metadata extraction works i

    Read More

    What Is Early Case Assessment (ECA)?

    Early case assessment (ECA) surfaces a matter's facts, risk, and cost early. Learn how ECA differs from review and guides the settle or litigate call.

    Read More

    Protect your evidence before spoliation becomes a problem.

    See how Venio Legal Hold helps your team issue, track, and document defensible holds in minutes.

    No credit card required • Free product tour available