AI use case

AI document intelligence for unstructured forms and documents

AI that takes documents in any format, such as scanned forms, PDFs, photos, emails and handwritten notes, splits and classifies them, extracts the required fields with a confidence score, validates them against business rules and source systems, and sends only the uncertain cases to a person before the data enters the downstream process.

By Len Debets · Last verified 27 September 2026 · 5 public deployments

At least 90%
Reported accuracy
Ancine, vendor claim.
10x
Reported productivity gain
Ancine, vendor claim.
USD 500,000 to USD 2.6 million
Indicative value per year
An organization that processes 500,000 forms and supporting documents a year. Worked example, see how it is calculated.

What problem does it solve?

Many organizations still run on documents that were designed for people: application forms, claims, certificates, supporting evidence, delivery notes, tax documents and letters, arriving as scans, phone photos, PDFs and email attachments. Staff open each one, work out what it is, retype the fields into a system and check them against other records. It is slow, error prone and hard to scale when volumes spike, and the backlog delays decisions that matter to citizens and customers.

Earlier OCR and template tools worked well for fixed layouts and struggled with anything else. Current document AI combines layout aware extraction, handwriting recognition and language models, so it can handle varied layouts, stamps, handwritten notes, tables across pages and several languages, as the Volvo Group deployment on this page shows. The design question is no longer whether AI can read the document but where it may act alone: straight through processing for confident, validated extractions, and human review for the rest, with every value traceable to the place on the page it came from.

How does it work?

  1. Ingest from every channel. Uploads, scans, email attachments and portal submissions land in one intake, where images are cleaned, rotated and split.
  2. Classify and separate. Each page or bundle is classified by document type (form, identity document, certificate, statement, invoice) and split into individual documents.
  3. Extract with confidence. Fields, tables, checkboxes and signatures are extracted, with the location on the page and a confidence score for each value; text in other languages can be translated.
  4. Validate. Values are checked against business rules (formats, totals, dates) and against source systems (the customer, case or supplier record).
  5. Route by confidence. Confident, valid documents flow straight into the downstream system; the rest go to a reviewer who sees the page and the extracted value side by side.
  6. Learn from corrections. Reviewer corrections are logged to improve extraction and to show which document types or sources cause errors.
Audience
Back office
Autonomy
Supervised agent
Adoption
Mainstream
Channels
API and system to system, Email, Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI document intelligence for unstructured forms and documents
KPIMedianReported rangeData pointsClaimed by
AccuracyToo few to pool
at least 90%
11 vendor
Hours savedNot pooled
10,000 hours
11 vendor
Productivity gainToo few to pool
10x
11 vendor

Value drivers: Lower cost to serve, Speed and cycle time, Employee productivity, Risk and loss reduction.

Indicative value

An organization that processes 500,000 forms and supporting documents a year

USD 500,000 to USD 2.6 million

Manual document handling cost avoided per year

How this is calculated

Formula: documents * minutesPerDocument * timeSaved * costPerMinute. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Documents processed per year documents, documents per year500,000500,000The reference organization.
Manual handling time per document today minutesPerDocument, minutes per document48Editorial assumption for classifying, keying and checking a document. Google Cloud reports that Pupuk Indonesia's data extraction took 5 to 10 minutes before AI.
Share of handling time removed, including review of uncertain cases timeSaved, fraction of handling time0.50.8Editorial assumption; review of low confidence documents stays with people.
Fully loaded processing staff cost costPerMinute, USD per minute0.50.8Editorial assumption, replace with your own.

What it leaves out: Counts only handling time. It leaves out faster decisions for customers and citizens, fewer keying errors, the cost of the platform and integration, and the reviewer capacity needed at peaks.

Who already uses it?

5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

U.S. Citizenship and Immigration Services

United States · Government and public sector · 2024

ProductionGrade B

Before this system, every page of an I-539 application (a request to extend or change nonimmigrant status) was scanned and stored as one document, which slowed adjudication and did not meet National Archives records standards. USCIS now uses an intelligent document processing tool to identify, classify and split each application into its component documents, such as the form itself, other USCIS forms, passports, driving licences, marriage certificates and bank statements. Pages the tool cannot identify go to a person. It is listed as deployed since November 2024; no outcome figures are published.

No outcome disclosed.

U.S. Immigration and Customs Enforcement

United States · Government and public sector · 2019

ProductionGrade B

Business units at ICE, part of the Department of Homeland Security, use an intelligent document processing platform (UiPath Suite and Azure AI Document Intelligence) with OCR and machine learning models to verify, extract and classify information from forms, automating repeatable work such as invoice processing and form entry validation. The agency lists it in operation since 2019 and says it saves staff significant time while improving data quality. No figures are published.

No outcome disclosed.

Ancine

Brazil · Government and public sector · 2026

ProductionGrade C

Ancine, which Google Cloud describes as the Brazilian cinema industry regulator, uses Google Cloud AI to extract and structure data from digitized tax documents to automate the accountability analysis of subsidized projects. Google Cloud reports extraction accuracy above 90% and a tenfold increase in analysts' daily processing capacity.

  • Accuracy: at least 90%, data extraction accuracy
    "This AI implementation achieved over 90% data extraction accuracy, boosting analysts' daily processing capacity by 10x."
    Claimed by: vendor
  • Productivity gain: 10x, analysts' daily processing capacity
    "This AI implementation achieved over 90% data extraction accuracy, boosting analysts' daily processing capacity by 10x."
    Claimed by: vendor

Pupuk Indonesia

Indonesia · Manufacturing · 2025

ProductionGrade C

Pupuk Indonesia, which Google Cloud describes as Asia's largest fertilizer producer, automated its document processing workflows with Vision AI and Gemini, working with Devoteam. Google Cloud reports that data extraction time fell from 5 to 10 minutes to 40 to 70 seconds, and that one employee now validates the results.

No outcome disclosed.

Volvo Group

Sweden · Automotive · 2023

ProductionGrade C

Volvo Group's automation team built a document processing solution on Azure AI Document Intelligence for its service and financing businesses. It reads emails, digital and scanned PDFs and written bills, including stamps, photographs, handwritten notes over printed text and tables across pages, translates content between languages and outputs XML or CSV for the receiving division, with processing time and success rate tracked in a dashboard. Microsoft reports that the solution has saved 10,000 manual hours since launch, about 850 hours a month.

  • Hours saved: 10,000 hours, since launch, about 850 hours per month
    "Since launch, the company has saved 10,000 manual hours—about 850-plus manual hours per month."
    Claimed by: vendor

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • A catalogue of document types with volumes and the fields each process needs
  • A labelled sample per document type to measure field level accuracy
  • Business rules and reference data for validation
  • Records rules for how originals and extracted data are kept

Systems to integrate

  • Scanning, mailroom, email and portal intake
  • Case management, ERP or line of business systems that receive the data
  • Reference data and customer or supplier master data for validation
  • Records and content management for the originals

Complexity: Medium

Extraction from common document types works out of the box. The effort goes into the long tail of layouts and poor scans, validation against source systems, confidence thresholds per field and a review interface that staff can work in quickly.

  1. 1

    Start with volume and pain

    Pick the document types with the highest volume and clearest fields, and measure today's handling time and error rate so the baseline is real.

  2. 2

    Measure accuracy per field

    Build a labelled test set per document type and measure accuracy per field, not per document. An average field accuracy of 95% can hide a date field that is wrong half the time.

  3. 3

    Set confidence thresholds per field

    Decide per field what confidence and which validation checks allow straight through processing, and start conservatively with more human review.

  4. 4

    Design the review screen

    Show the page region next to each extracted value and let reviewers correct with one click. Review speed decides most of the business case.

  5. 5

    Validate against systems of record

    Check names, numbers and totals against the case, customer or supplier record before data is accepted, and flag mismatches rather than overwrite.

  6. 6

    Watch for drift

    Track corrections by document type and source, and retest when forms, suppliers or scanning change.

Guardrails

  • Straight through processing only above field level confidence thresholds and after validation checks
  • Every extracted value linked to its location in the source document
  • Unrecognized pages and documents always go to a person
  • Personal data in documents processed and stored under the same controls as the source system
  • Extraction informs decisions; eligibility, benefit or credit decisions stay with the owning process and people

KPIs to instrument

  • Field level accuracy per document type on a labelled sample
  • Straight through processing rate per document type
  • Reviewer time per document and correction rate
  • End to end time from receipt to data available in the downstream system
  • Errors found downstream that originated in extraction

Human in the loop

Reviewers handle every document below the confidence threshold or failing validation, and their corrections are logged. Process owners set thresholds and approve changes to them, and quality teams sample straight through documents regularly to confirm accuracy holds.

Common failure modes

High average, weak critical field
Overall accuracy looks good while one field that drives decisions is often wrong. Measure and threshold per field.
Silent errors in straight through processing
Confident but wrong values enter systems unchecked. Validate against source systems and sample straight through documents.
The long tail stalls the program
Rare layouts consume the project. Route them to people and automate by volume.
Extraction becomes the decision
A missing field triggers an automatic rejection of an application. Keep decisions in the owning process with human review.

What are the risks and rules?

EU AI Act

Depends on design

Classifying documents and extracting data for a person or process to use is usually minimal risk. Even inside an Annex III area, a system that only performs a narrow procedural task, such as splitting and classifying documents, can fall outside the high risk category under Article 6(3); the provider must document that assessment and register the system (Article 6(4) and Article 49(2)). The picture changes when extraction materially influences decisions in Annex III areas, such as eligibility for public assistance benefits (point 5(a)), creditworthiness (point 5(b)) or asylum, visa and residence permit applications (point 7), where the whole system must be assessed as potentially high risk. The Article 6(3) exception never applies when the system performs profiling of natural persons.

Guidance

Controls to put in place

  • Documented accuracy per field and document type before go live and after changes
  • Threshold and routing rules under change control
  • Audit trail from each extracted value to the source document and reviewer action
  • Retention of originals and extracted data aligned with records rules
  • Assessment of whether downstream decisions fall in an Annex III area

Frequently asked questions

How accurate is AI document extraction?
It depends on the document type and the field, so measure it per field on your own documents. Google Cloud reports that Ancine, Brazil's cinema industry regulator, reached over 90% data extraction accuracy on digitized tax documents and a tenfold increase in analysts' daily processing capacity. Set confidence thresholds per field and keep people on the uncertain cases.
What volumes and savings do organizations report?
Microsoft reports that Volvo Group's solution for invoices, credit notes and claims documents has saved 10,000 manual hours since launch, about 850 a month. Google Cloud reports that Pupuk Indonesia cut data extraction time from 5 to 10 minutes to 40 to 70 seconds, with one employee validating the results. USCIS splits and classifies I-539 applications so that adjudicators find each supporting document faster; it publishes no figures.
How is this different from invoice processing or correspondence triage?
Invoice processing is one specialized use of document AI, with purchase order matching and posting. Correspondence triage is about routing incoming mail. This page covers the general capability for forms, applications and supporting documents in any process.

How to cite this page

Blits.ai AI Use Case Library, "AI document intelligence for unstructured forms and documents", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/intelligent-document-processing. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryBanking

AI for supplier invoice processing in accounts payable

AI that captures supplier invoices from any format, extracts header and line data, matches them to purchase orders and goods receipts, proposes tax and cost centre coding, flags duplicates and suspected fraud, and routes them for approval and posting, leaving only exceptions to accounts payable staff.

Deployments
4 public, best grade B
Reported productivity gain
80%
Kingfisher, organization claim
Cross industryBanking

AI for inbound correspondence triage and routing

AI that sorts inbound correspondence before anyone answers it: it takes every inbound letter, email, upload and secure message into one intake, identifies what it is, extracts the key fields, links it to the right customer and account, sets priority and routes it to the right team or workflow, replacing the manual sorting desk.

Deployments
6 public, best grade B
Reported accuracy
91%
Travelers, vendor claim
Insurance

AI for commercial underwriting submission intake and triage

AI that reads incoming broker submissions for commercial insurance (emails, applications, schedules of values, loss runs and supplements), extracts the risk data into a structured record, checks clearance and appetite, enriches the risk with internal and third party data and ranks it, so underwriters open a complete, prioritized file instead of an inbox.

Deployments
9 public, best grade B
Reported accuracy
about 98%
Paragon Insurance Group, organization claim
BankingPayments and cards

AI for application and identity fraud detection

AI that checks incoming account and loan applications for forged or AI generated documents, synthetic and stolen identities, and coordinated application rings, by analysing documents, device and application data across the whole queue and cross checking against bureau and official sources.

Deployments
6 public, best grade B
Reported detection improvement
2.5x
Department for Work and Pensions, organization claim
Banking

AI examination of trade documents under letters of credit and collections

AI that reads the full document presentation under a letter of credit or collection (bill of lading, commercial invoice, packing list, certificates), extracts and cross checks the data, tests it against the instructions and the ICC rules (for letters of credit, the credit terms, UCP 600 and ISBP), and lists discrepancies by severity with the rule cited, so qualified examiners focus on the genuine exceptions.

Deployments
3 public, best grade B
Autonomy
Supervised agent
BankingInsurance

AI for back office account servicing execution

AI that executes the servicing requests that land in operations queues, such as address and mandate changes, standing instructions, beneficiary updates, reissues, payoff and reference letters and loan maintenance, by reading the request, checking it against policy and entitlements, and preparing or making the change in core systems under dual control.

Deployments
2 public, best grade C
Reported cycle time reduction
95%
SS&C Technologies, vendor claim