What problem does it solve?
Travel and expense spend is high volume, low value per transaction, and hard to police at scale. Traditional audit teams cannot review every report line by line, so they sample: only expenses above a value threshold, or a random percentage, get a real look. At Takeda, AppZen reports that at times 70 to 100% of expenses triggered audit rules, yet sampling and value thresholds still left spend uncovered and exposed to duplicate submissions and employee spend leakage that a person reviewing one report at a time cannot see across the whole population.
Multinational organizations add language and currency variation across dozens of countries, each with its own per diem and travel policy, on top of the volume problem. The result is a familiar trade off: either spend more headcount on audit, which does not scale with travel volume, or accept that most spend goes unchecked and hope the exceptions surface some other way, for example when a manager happens to notice.
How does it work?
- Capture every report. Expense reports, receipts and the underlying corporate card transactions feed into one pipeline from the T&E platform, so every line has a receipt image or card record to check against, not just the ones a person opens.
- Score every line against policy. Each line is checked against the organization's written policy by category, region and grade, rather than a single value threshold, and against the employee's own submission history.
- Cross check for duplicates and altered receipts. The agent compares receipts across reports and employees for duplicate submissions (including the same meal claimed by two attendees) and checks receipt images for signs of alteration.
- Auto approve the clean majority. Lines that pass every check within policy limits approve automatically; anything flagged goes to an auditor with the receipt, the rule it triggered and similar past decisions shown together.
- Learn from auditor decisions. Confirmed and overturned flags feed back as candidate rule adjustments, which a T&E or finance manager reviews and approves before they change what auto approves.
- Feed the policy owners. Patterns by category, region or employee go to the corporate card and travel policy teams, so recurring issues get fixed at the policy level, not flagged again every month.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Early adopters
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Automation rate | Too few to pool | 63% to 72% | 2 | 2 vendor |
| Hours saved | Not pooled | 3000 hours to 4000 hours | 2 | 2 vendor |
| Interactions handled | Not pooled | about 400,000 | 1 | 1 vendor |
Value drivers: Lower cost to serve, Risk and loss reduction, Compliance quality, Employee productivity.
Indicative value
A company with 7,000+ employees submitting about 130,000 expense reports a year
USD 64,838 to USD 133,380
Manual expense audit effort avoided per year
How this is calculated
Formula: reportsPerYear * autoApprovalShare * minutesPerReport / 60 * costPerHour. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Expense reports submitted per year reportsPerYear, reports per year | 130,000 | 130,000 | AppZen reports that before AppZen, Databricks' "two auditors manually reviewed nearly 130K expense reports annually" across "7,000+ employees globally." Replace with your own report volume. Source |
| Share of reports the AI approves without a person autoApprovalShare, fraction of reports | 0.63 | 0.72 | Range spans the auto approval rates AppZen reports for its two customers: 63% at [Takeda](https://www.appzen.com/resources/case-studies/how-takeda-is-transforming-global-expense-auditing-with-ai) across 63 countries, 72% at Databricks, where AppZen says the team kept adjusting its configuration month after month. Expect a lower rate before your own policy is tuned. Source |
| Auditor minutes saved per auto approved report minutesPerReport, minutes per auto approved report | 1.9 | 1.9 | Derived from the hours AppZen reports Databricks saved, spread only across the reports that were auto approved: about 1.9 minutes per auto approved report (nearly 3,000 auditor hours saved a year across 130,000 reports at a 72% auto approval rate), which anchors this figure to the reference org above. This is time saved on the reports the AI clears, not a manual review time per report; no source gives that figure. AppZen also reports [Takeda](https://www.appzen.com/resources/case-studies/how-takeda-is-transforming-global-expense-auditing-with-ai) saving 4,000 auditor hours a quarter across "400K expense audits annually" at a 63% auto approval rate, but that figure is not used here: AppZen never states that an expense audit is the same unit as an expense report, so it cannot be safely converted into minutes per report and applied to the reference org's report count. Replace with your own time study. Source |
| Fully loaded cost of a T&E auditor costPerHour, USD per hour | 25 | 45 | Editorial assumption for a blended onshore and offshore audit team. |
What it leaves out: Labour only. It leaves out the wasteful and non compliant spend actually caught (AppZen reports Databricks identified $483K in wasteful spend over twelve months), faster employee reimbursement, and the cost of the platform and policy configuration work.
Who already uses it?
2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Databricks
United States · Technology and software · 2025
Before AppZen, Databricks' four person global audit team relied on manager approval, which gave "no visibility, no forensics, and no data analysis on receipts," plus a four eyes check where two auditors manually reviewed nearly 130,000 expense reports a year across more than 7,000 employees. AppZen's Expense Audit, layered onto the existing Emburse Chrome River expense system, automated that review for duplicates and policy risk. Custom AppStore models such as Double-Dip Detection catch employees who submit a meal expense while also being listed as an attendee on someone else's expense, something the team could not catch before AppZen.
- Hours saved: about 3000 hours, per year
"In one year, Databricks saved nearly 3,000 manual auditor work hours."
Claimed by: vendor - Automation rate: 72%
"Databricks identified $483K in wasteful spend and saved 3,000 auditor hours annually with AppZen, achieving 72% auto-approval for 9,000 employees."
Claimed by: vendor
Takeda
Japan · Pharma and life sciences · 2022
Takeda deployed AppZen's Expense Audit solution to review 100% of employee expense reports across its global operations, replacing a manual process that, at times, flagged as much as 70 to 100% of expenses for audit based on trigger rules yet still left the company vulnerable to duplicates and employee spend leakage. The AI models include translation capability to handle the language differences across Takeda's European and Asian operations, and auditors now focus their attention on high risk expenses.
- Automation rate: 63%
"Takeda achieved 63% auto-approval across 63 countries, processing 400K expense audits annually while saving 4,000 auditor hours quarterly with AppZen."
Claimed by: vendor - Interactions handled: about 400,000, per year
"Takeda achieved 63% auto-approval across 63 countries, processing 400K expense audits annually while saving 4,000 auditor hours quarterly with AppZen."
Claimed by: vendor - Hours saved: about 4000 hours, per quarter
"Takeda achieved 63% auto-approval across 63 countries, processing 400K expense audits annually while saving 4,000 auditor hours quarterly with AppZen."
Claimed by: vendor
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Twelve months of history of audited expense reports with the outcome and reason
- The written expense policy by category, region and grade, in a machine readable form
- Card and travel booking feeds to cross check receipts against actual charges
Systems to integrate
- T&E and expense platform (SAP Concur, Expensify, Emburse, Ramp)
- Corporate card and travel booking feeds
- HR system for employee grade, location and manager
- ERP or payroll for reimbursement posting
Complexity: Medium
Checking a receipt against a rule is straightforward; the work is encoding a policy that differs by category, region and grade into machine readable rules, and tuning thresholds so genuine risk is not buried under false positives on routine claims like meals and mileage.
- 1
Baseline current coverage and findings
Record what share of reports get a real manual look today, what the sampling rule is, and what the audit team actually finds, so the AI is credited only for the increment.
- 2
Encode the policy as machine readable rules
Turn the written expense policy into structured rules by category, region and grade; most early gaps come from policy that only exists as prose no one reads consistently.
- 3
Start with the highest risk categories
Pick two or three categories with the most manual audit volume or the most prior findings, such as meals, mileage or entertainment, rather than trying to cover every category at once.
- 4
Run in parallel with current sampling
Let the AI score every report next to the existing sample based process for a full close cycle, and compare what each approach would have flagged before changing anything live.
- 5
Automate approval within limits
Auto approve only lines that pass every check within an agreed value limit; anything flagged, anything above the limit, and anything the model has not seen before goes to a person.
- 6
Feed the policy and card teams
Route recurring flag patterns by category or employee to the people who own travel policy and the corporate card program, so the root cause gets fixed, not just the individual claim.
Guardrails
- Every report checked against policy, not a value threshold sample; 100% coverage is the point
- Only clean, policy compliant lines under an agreed value limit auto approve; anything flagged goes to a human auditor
- Duplicate and altered receipt checks run on every submission before reimbursement
- Any flag that could support disciplinary or employment action goes to a manager or HR review; the AI never decides or triggers that action itself
KPIs to instrument
- Auto approval rate by category and region
- Auditor hours per period
- Wasteful or non compliant spend identified
- Reimbursement cycle time
- Repeat flag rate by employee and category
Human in the loop
Auditors review every flagged line and decide whether to reject it, ask the employee for more information, or approve it. A T&E or finance manager approves changes to policy rules and thresholds, and reviews a sample of auto approved lines every month for drift.
Common failure modes
- False positives bury real risk
- Too many low value flags on routine claims train auditors to rubber stamp the queue. Tune thresholds by category and track the override rate, not only the flag count.
- New patterns the model has not seen
- A scheme built around the model's blind spot, such as a generated receipt image, goes through. Sample auto approved lines and periodically red team the checks with new patterns.
- Flags treated as verdicts
- A flag is acted on as if it were a finding rather than a lead, which is unfair to the employee if the flag is wrong. Require a human decision and a documented reason before any action follows from a flag.
What are the risks and rules?
EU AI Act
High risk
Annex III point 4(b) covers AI systems intended to monitor and evaluate the performance and behaviour of persons in a work related relationship. Scoring every line against the employee's own submission history, and instrumenting a repeat flag rate by employee, is that kind of behavioural evaluation, so this design falls under Annex III. The only carve out, Article 6(3), lets a narrow procedural or preparatory task escape high risk with a documented assessment, but the last subparagraph of Article 6(3) removes that carve out whenever the system performs profiling of natural persons. Scoring lines against an individual employee's history is profiling, so the carve out is not available here: keeping a human auditor as the actual decision maker on any personnel action is a required control, not an exit from Annex III.
Guidance
- Annex III, point 4(b): employment, workers' management and access to self employment (European Union, Europe). Lists AI systems intended to make decisions affecting the terms of a work related relationship, its promotion or termination, to allocate tasks based on individual behaviour or personal traits, or to monitor and evaluate the performance and behaviour of persons in that relationship, as high risk.
- Article 6(3): the narrow task exception, and why it does not apply here (European Union, Europe). An Annex III system escapes high risk only if it performs a narrow procedural task, improves the result of a previously completed human activity, detects deviations from prior human decision making patterns without replacing or influencing the completed human assessment, or is preparatory. The last subparagraph closes this exception whenever the system profiles natural persons, which per employee, history based scoring does.
Controls to put in place
- Every flag reviewed by a human auditor before it affects reimbursement or any personnel action
- Documented, versioned policy rules with an owner and a review date
- Consistency sampling across regions, expense categories and employee grades to catch uneven flagging
- Full audit trail of every flag, the rule it triggered, and the human decision that followed
Frequently asked questions
- What share of expense reports can AI approve without a person?
- AppZen reports a 63% auto approval rate across 63 countries at Takeda and a 72% auto approval rate at Databricks, where AppZen says the team kept adjusting its configuration month after month. Expect a lower rate at first, since the policy rules and thresholds need tuning against your own spend patterns before the model earns a wider auto approval limit.
- Does full coverage replace sampling?
- It changes what sampling is for. Instead of choosing which reports get a real look, every report is checked against policy and prior patterns, and the sample becomes a quality check on the AI itself: auditors periodically review a slice of auto approved lines to catch drift, not a slice of all submissions to find the risky ones.
- Is expense audit AI high risk under the EU AI Act?
- As designed here, yes. Annex III point 4(b) covers AI systems that monitor and evaluate employee behaviour, and scoring lines against an employee's own history is that kind of evaluation. The Article 6(3) narrow task exception cannot rescue it, because that exception never applies once a system profiles individual people, which per employee, history based scoring does. Keeping a human auditor as the actual decision maker on any personnel action is a required control under Annex III, not a way around it.
- What should stay with a person?
- Any decision with disciplinary or termination consequences, ambiguous policy interpretation that a rule cannot capture, and cross border cases with tax or immigration implications.
How to cite this page
Blits.ai AI Use Case Library, "AI agent for travel and expense report audit", last verified 28 September 2026, https://www.blits.ai/ai-use-cases/travel-and-expense-audit-agent. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 28 September 2026: First published