AI use case

AI for eDiscovery and disclosure document review

AI that sorts, prioritises and codes large collections of emails, chats and files for relevance, issues and legal privilege in litigation, investigations and regulatory requests, so that lawyers review the documents most likely to matter and can show the court how the rest were handled.

By Len Debets · Last verified 27 September 2026 · 4 public deployments

85%
Reported cycle time reduction
Purpose Legal, vendor claim.
About 30 million
Interactions handled
Serious Fraud Office (organization claim).
USD 225,000 to USD 1.5 million
Indicative value per year
A litigation or investigations team that reviews 1 million documents a year. Worked example, see how it is calculated.

What problem does it solve?

Every lawsuit, investigation and regulatory information request starts with a collection of electronic material that nobody can read in full. A single civil matter can involve more than 300,000 documents, and some Serious Fraud Office cases start with up to 60 million. Lawyers still have to find what is relevant, what supports or undermines each side, and what is privileged, under deadlines set by a court or a regulator.

Keyword searches were the first answer, and they are weak: they miss documents that use other words, return large volumes of noise, and give a false sense of completeness. Linear review by contract lawyers is slow and expensive, and the US Department of Justice lists errors and speed delays in exclusively manual review of voluminous electronic information as the problem its eLitigation tools address. In criminal cases the stakes are higher still. The UK Serious Fraud Office offered no evidence in its G4S case after ten years, with disclosure featuring as a core reason, and its inspectorate found that a misunderstanding of how searches worked in its older review system compounded the problems.

How does it work?

  1. Collect and process. Material from mailboxes, devices, chat tools and file shares is loaded into a review platform, text is extracted, duplicates and near duplicates are removed and email threads are grouped.
  2. Explore early. Clustering, timelines and concept search show what the collection is about, so the team can exclude irrelevant data and agree search parameters before paid review starts.
  3. Rank or code. Either a classifier learns from reviewers' decisions and keeps ranking the remaining documents by likely relevance (active learning), or a large language model reads each document against written review instructions and returns a relevance call, the issues it touches and a short rationale.
  4. Screen for privilege and sensitivity. A separate pass flags likely privileged documents, personal data and material that needs redaction, for lawyers to confirm.
  5. Validate. Random samples from the documents classed as relevant and not relevant are reviewed by people to estimate precision, recall and elusion, and the method and results are recorded so they can be explained to the other side or the court.
  6. Review and produce. Lawyers review the prioritised set, decide what to disclose, withhold or redact, and the platform keeps the audit trail.
Audience
Employee facing
Autonomy
Copilot
Adoption
Mainstream
Channels
Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI for eDiscovery and disclosure document review
KPIMedianReported rangeData pointsClaimed by
Interactions handledNot pooled
300,000 to 30 million
21 organization, 1 vendor
Cost savingsNot pooled
at least USD 70,000
11 vendor
Cycle time reductionToo few to pool
85%
11 vendor
Hours savedNot pooled
4000 hours
11 vendor

Value drivers: Lower cost to serve, Speed and cycle time, Employee productivity, Compliance quality, Risk and loss reduction.

Indicative value

A litigation or investigations team that reviews 1 million documents a year

USD 225,000 to USD 1.5 million

First pass review cost avoided per year

How this is calculated

Formula: documents / 1000 * hoursPerThousand * reviewAvoided * hourlyCost. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Documents collected for review per year after deduplication documents, documents per year1,000,0001,000,000The reference team. Replace with your own review volumes.
Reviewer hours per 1,000 documents in a linear first pass review hoursPerThousand, hours per 1,000 documents1525Editorial assumption, replace with your own review rates.
Share of first pass review hours avoided by prioritisation or AI coding reviewAvoided, fraction of review hours0.30.6Conservative against the vendor reported 85% reduction in review time for one Purpose Legal matter on this page, because validation sampling, privilege review and quality control still need people.
Fully loaded cost of a reviewer hour hourlyCost, USD per hour50100Editorial assumption covering contract reviewers and supervising associates.

What it leaves out: Covers first pass review labour only. It leaves out platform and model costs, processing and hosting fees, the cost of validation and of any dispute about the method, and the value of meeting deadlines that manual review would miss.

Who already uses it?

4 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

U.S. Department of Justice

United States · Government and public sector · 2024

ScaledGrade B

Department of Justice components use commercial eLitigation platforms such as Relativity and Everlaw for investigations, litigation and FOIA and Privacy Act work, and these tools increasingly include AI. The department's inventory entry describes surfacing potentially discoverable material in large collections of emails, text messages and other records, locating potentially inculpatory or exculpatory evidence, and identifying material to disclose or withhold under legal rules and privileges. The entry is department wide and deployed since January 2024. It is formally classified as "Presumed High-Impact, but Not High-impact", because the output is not the principal basis for decisions with legal or similarly significant effect, while noting that the tools are used in high impact contexts and that individual uses vary. No outcome figures are published.

No outcome disclosed.

Federal Trade Commission

United States · Government and public sector · 2020

ProductionGrade B

The Federal Trade Commission's Bureau of Competition, Bureau of Consumer Protection and Office of the General Counsel use Relativity Active Learning in eDiscovery for consumer protection and competition investigations and litigation. The inventory entry gives the problem as manual document coding being very time consuming during legal review and the output as predicted pertinent documents. It lists the use as deployed since 2020 and not high impact. No outcome figures are published.

No outcome disclosed.

Serious Fraud Office

United Kingdom · Government and public sector · 2018

ScaledGrade B

The UK Serious Fraud Office first used an AI tool in its Rolls-Royce investigation to screen about 30 million documents for material potentially covered by legal professional privilege, work that independent barristers had previously done by hand. From April 2018 it made AI document review available to all new cases and adopted OpenText Axcelerate, which groups documents by subject, builds timelines and removes duplicates; at launch the SFO said it would eventually be able to sift for relevancy. The 2024 inspection by HM Crown Prosecution Service Inspectorate describes Axcelerate's artificial intelligence and machine learning features and how document reviewers tag material for relevancy in it. It warned that there remains a risk that some staff are not confident using the system, found that a number of staff were not using it to its full potential and that many viewed its training as inadequate, and warned that searching millions of documents is not an exact science. It also found that disclosure problems in an earlier case were compounded by a misunderstanding of how searches worked in the SFO's previous document review system, which is no longer used.

  • Interactions handled: about 30 million, Rolls-Royce investigation pilot, privilege screening
    "It enabled the estimated 30 million documents provided by the company to be analysed for material potentially covered by Legal Professional Privilege."
    Claimed by: organization

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Processed collections with extracted text, metadata and deduplication
  • Written review instructions per issue, with examples of relevant and not relevant documents
  • A seed or control set coded by lawyers who know the case
  • A sampling plan for validation with target recall agreed in advance

Systems to integrate

  • eDiscovery or review platform (processing, review, production)
  • Legal hold and collection tools for mailboxes, devices and chat
  • Matter management and privilege log tooling
  • Secure hosting in the jurisdiction the data must stay in

Complexity: Medium

The technology is mature and widely available inside review platforms. The work is in the review protocol, the validation method, privilege handling and agreeing the approach with the other side, the court or the regulator, and in training case teams to use it properly.

  1. 1

    Agree the protocol before the model

    Write the review questions, issue definitions and privilege rules first, and decide how success will be measured (recall target, sample sizes). Where the other side or a court will scrutinise the process, share the approach early rather than defend it later.

  2. 2

    Pilot on a coded sample

    Run the classifier or the language model on a few hundred documents that lawyers have already coded, compare, and refine the instructions until the disagreements are understood. Relativity reports that Purpose Legal reached a workable prompt for ten issues after three iterations on a sample of fewer than 500 documents.

  3. 3

    Run, rank and route

    Apply the model to the full population, send the highest ranked documents to human review first, and keep a separate privilege and personal data pass.

  4. 4

    Validate with statistics, not impressions

    Draw random samples from both the relevant and the not relevant sets, have people review them blind, estimate recall and elusion, and document every step so it can be explained.

  5. 5

    Train the case team

    Make sure the people running the review understand what the tool does and does not do. The inspectorate warned of a risk that some SFO staff are not confident using the new platform, found that a number were not using it to its full potential and that many saw the training as inadequate, and found that misunderstandings about search had contributed to earlier failures.

Guardrails

  • No document is withheld as privileged or produced without a lawyer's decision
  • Validation sampling with a recall estimate before any review is declared complete
  • Written record of search terms, model instructions, versions and sampling results
  • Separate handling of privileged and personal data, with redaction checked by a person
  • Data stays in the hosting region and is not used to train a vendor's general models

KPIs to instrument

  • Estimated recall and elusion rate from validation samples, per matter
  • Reviewer hours per 1,000 documents, before and after
  • Share of the collection reviewed by people
  • Days from collection to production
  • Privilege clawback requests and disclosure errors found later

Human in the loop

Lawyers write the review instructions, code the training or validation samples, decide every privilege call and every production, and sign off the validation results. The AI decides only the order and the first view of relevance.

Common failure modes

False confidence in completeness
Teams treat search or model output as if it found everything. HMCPSI warned that searching millions of documents is not an exact science. Measure recall and say what it is.
Processing gaps upstream
Documents that were never extracted or indexed properly cannot be found by any model. Check processing exceptions, container files and encoding before trusting results.
Instructions that drift from the case
The issues change as the case develops but the model instructions do not. Version the instructions and rerun validation when they change.
Unexplainable method
The team cannot describe how documents were excluded when challenged. Keep the protocol, versions and samples as part of the matter record.

What are the risks and rules?

EU AI Act

Depends on design

Document review for a party in civil litigation or an internal investigation is not listed in Annex III, so it is usually minimal risk. It becomes high risk where a law enforcement authority uses AI to evaluate the reliability of evidence in the investigation or prosecution of criminal offences (Annex III point 6(c)), or where a judicial authority uses it to research and interpret facts and law (point 8(a)). Prosecutors and investigators should classify each use against those points.

Guidance

Controls to put in place

  • Documented review protocol and validation plan for every matter
  • Named lawyer accountable for the review method and its explanation
  • Audit trail of model versions, instructions, coding decisions and samples
  • Data processing agreements and hosting location that match the data's origin
  • Quality assurance batches reviewed by a second person

When it went wrong elsewhere

  • SFO drops decade-long probe and prosecution and announces further issue with legacy disclosure system. The Law Gazette reports that the Serious Fraud Office found a further problem in its legacy Autonomy review system, in how some digital container files were expanded, so that some items may not have been available for review in about 20 long running cases; an earlier review covered 66 historical conviction cases. Not an AI model error, but a reminder that review results are only as complete as the processing underneath them.

Frequently asked questions

Can we rely on AI review when the other side or the court will scrutinise it?
Only with a method you can explain. In England and Wales the Attorney General's disclosure guidelines let investigators use search techniques when reviewing everything would be disproportionate, provided they record their reasons, and the judiciary is reviewing civil disclosure rules with technology assisted review and AI in mind. Agreed instructions, statistical validation and a written record are what make the result defensible.
How much review time does it save?
It depends on the collection and the recall target. Relativity reports that Purpose Legal cut project time by 85% (4,000 review hours) on one 300,000 document matter, measured against a contract review that would have taken multiple weeks and was never run. The UK Serious Fraud Office used its pilot tool to screen about 30 million documents for privilege in the Rolls-Royce case, work that independent barristers had done by hand. Plan on people still reviewing the prioritised set, the privilege calls and the validation samples.
What is the difference between active learning and generative AI review?
Active learning trains a classifier on reviewers' decisions and keeps reranking the rest; the FTC uses Relativity Active Learning to predict pertinent documents. Generative AI review has a language model read each document against written instructions and explain its call. Both need the same validation.
Is AI document review high risk under the EU AI Act?
For a company or law firm reviewing documents in civil litigation, usually not. It is high risk when law enforcement uses it to evaluate the reliability of evidence in criminal cases, or when a judicial authority uses it to research and apply the law.

How to cite this page

Blits.ai AI Use Case Library, "AI for eDiscovery and disclosure document review", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/ediscovery-and-disclosure-document-review. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Government and public sector

AI for court and case file summarization

AI that condenses court filings, case files, evidence recordings and earlier decisions into structured summaries, chronologies and draft case reports with references to the source pages, so that judges, prosecutors, tribunal staff and government lawyers find what matters faster, while the person responsible reads the underlying material and makes every legal judgment.

Deployments
4 public, best grade B
Autonomy
Copilot
Government and public sector

AI for freedom of information request processing

AI that helps a public body handle freedom of information and open government requests: logging and clarifying requests, spotting duplicates, searching and deduplicating the records in scope, proposing redactions with the exemption that applies, and drafting the response letter, with an FOI officer deciding what is released.

Deployments
4 public, best grade B
Autonomy
Copilot
Cross industryGovernment and public sector

AI document intelligence for unstructured forms and documents

AI that takes documents in any format, such as scanned forms, PDFs, photos, emails and handwritten notes, splits and classifies them, extracts the required fields with a confidence score, validates them against business rules and source systems, and sends only the uncertain cases to a person before the data enters the downstream process.

Deployments
5 public, best grade B
Reported accuracy
at least 90%
Ancine, vendor claim
Cross industryBanking

AI for inbound correspondence triage and routing

AI that sorts inbound correspondence before anyone answers it: it takes every inbound letter, email, upload and secure message into one intake, identifies what it is, extracts the key fields, links it to the right customer and account, sets priority and routes it to the right team or workflow, replacing the manual sorting desk.

Deployments
6 public, best grade B
Reported accuracy
91%
Travelers, vendor claim