What problem does it solve?
Tax administrations can examine only part of the returns they receive, so the question of which returns to open decides both how much revenue is protected and who carries the burden of an audit. Traditional selection relies on fixed rules, random samples and the judgment of experienced staff. That misses new schemes, sends officers to returns that turn out to be correct (so called no change audits) and struggles with complex taxpayers, such as large partnerships, where the risk is spread over many entities and schedules.
More data now supports model based selection: third party reporting, electronic invoicing, bank and cross border information exchange. It has also shown the risk. A selection model that is accurate on average can still concentrate audits on particular groups, and the Dutch childcare benefits scandal, in which the tax administration's risk classification used nationality, made fairness, transparency and human review preconditions rather than extras.
- The IRS projects an annual gross tax gap of USD 696 billion for tax year 2022, of which USD 539 billion comes from tax understated on timely filed returns.IRS: The tax gap (2024)
How does it work?
- Assemble the risk picture. Returns are joined with third party data the administration already holds: employer and bank reporting, invoices, customs data, prior audit results and information exchanged with other countries.
- Score. Anomaly detection flags returns that break a taxpayer's own pattern or deviate from peers; supervised models trained on past audit outcomes estimate the likelihood and size of a correction; business rules encode known risks. The approaches can be combined; the Belastingdienst VAT signal model on this page, for example, is purely rules based.
- Explain the signal. Each score comes with the features and rules that drove it, so an officer can see why a return was flagged and challenge it.
- Select with people. Risk teams turn scores into case lists, mixed with a random sample that keeps measuring the unflagged population, and an officer decides whether to open a check.
- Learn and monitor. Audit results feed back into the model; selection rates and outcomes are compared across groups to detect disparate impact before it becomes a scandal.
- Audience
- Back office
- Autonomy
- Assist
- Adoption
- Early adopters
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Interactions handled | Not pooled | about 1500 | 1 | 1 organization |
| Users served | Not pooled | about 5500 | 1 | 1 organization |
Value drivers: Risk and loss reduction, Employee productivity, Compliance quality, Lower cost to serve.
Indicative value
A national tax administration that completes 10,000 desk and field audits a year
EUR 2 million to EUR 27 million
Additional tax assessed from the same audit capacity per year
How this is calculated
Formula: audits * yieldPerAudit * uplift. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Audits completed per year audits, audits per year | 8,000 | 12,000 | Editorial assumption. Replace with your own audit volume. |
| Average additional tax assessed per audit yieldPerAudit, EUR per audit | 5,000 | 15,000 | Editorial assumption. Replace with your own average yield, including audits that end with no change. |
| Relative increase in yield from better selection uplift, fraction of yield | 0.05 | 0.15 | Editorial assumption, deliberately modest. The public evidence on this page reports usage and selection practice, but no audited tax yield uplift. |
What it leaves out: Assessed tax, not collected tax. It leaves out collection losses, appeals, the cost of building and assuring the models, the deterrence effect and the benefit to compliant taxpayers of fewer no change audits.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
HM Revenue and Customs
United Kingdom · Government and public sector · 2025
HMRC's VAT Return Analysis Tool brings a VAT trader's entity, ledger and return data for the most recent seven years into one interactive view for VAT officers, and uses a classical statistical model (seasonal trend decomposition with an interquartile range rule) to flag anomalous values in the return history. Officers use it to prepare and carry out compliance checks; the tool makes no decisions, and any assessment is made by an officer and can be appealed through the normal route. HMRC's transparency record says around 5,500 officers are licensed and the tool is used about 1,500 times a day.
- Users served: about 5500, licensed VAT officers
"Around 5,500 officers have a license to use the tool as part of their VAT compliance work."
Claimed by: organization - Interactions handled: about 1500, per day
"There are ~1,500 daily uses of the tool."
Claimed by: organization
Belastingdienst
Netherlands · Government and public sector · 2024
Since 1 October 2024 the Dutch Tax and Customs Administration has supported staff of its Large Businesses directorate with a signal model that risk assesses VAT returns. The model applies business rules drawn from legislation, expertise and statistics, is explicitly not self learning, and sorts signals into priority categories that help staff decide when and by whom a signal is handled. The register says the model also takes decisions itself; where it cannot (more complex situations or a deviation in the return), a staff member intervenes. Samples are also drawn from returns the model does not select, and the rules are evaluated every year against those results. It is one of several VAT signal models the Belastingdienst has published in the national algorithm register.
No outcome disclosed.
Internal Revenue Service
United States · Government and public sector · 2023
In September 2023 the IRS announced that it was expanding its Large Partnership Compliance programme, which it described as a pilot leveraging AI, to additional large partnerships. It said the returns had been selected with the help of AI by data scientists and tax enforcement experts who applied machine learning to identify compliance risk in partnership tax, general income tax and accounting, and international tax. The IRS said it would open examinations of 75 of the largest partnerships, each with more than USD 10 billion in assets on average, a segment that had seen little examination coverage. The IRS also said AI would help improve case selection so that fewer taxpayers face audits that end with no change. No outcome of the AI selected examinations is published in the release.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Historical audit outcomes, including no change results, linked to returns
- Third party information returns and invoices, with a documented legal basis for each
- A random audit or sample programme to measure the unflagged population
Systems to integrate
- Return processing and taxpayer account systems
- Case management for compliance checks and audits
- Data warehouse with third party and exchange of information data
- Analytics workspace for risk teams and model monitoring
Complexity: High
The data usually exists; the difficulty is governance. Selection models touch taxpayer rights, need a legal basis for each data source, must be explainable to officers and courts, and need fairness monitoring that often lacks the protected characteristic data to do it well.
- 1
Start where the data is richest and the harm is lowest
Begin with business taxes such as VAT, where invoices and returns give a strong signal and the population is mostly companies, before models touch individuals and families.
- 2
Keep a random sample running
Reserve part of the audit capacity for random selection. It is the only way to measure what the model misses and to prove it beats the old approach.
- 3
Put explanation in front of the officer
Show the reasons for a flag next to the return. HMRC's VAT tool, for example, shows expected against observed values for each return period, so the officer judges the anomaly.
- 4
Test for disparate impact before and after go live
Compare selection rates and hit rates across groups you can measure, including proxies such as income band and region, and document how you will act on a gap.
- 5
Publish what you run
Register every selection model with its purpose, data and human oversight. The Dutch Belastingdienst publishes its selection and signal models in the national algorithm register.
Guardrails
- No protected characteristic, or obvious proxy such as nationality, as a model feature
- A score never triggers an assessment or penalty on its own; an officer decides every action
- Random sample alongside model selection to measure performance and fairness
- Documented legal basis for every data source used in scoring
- Model changes reviewed by an independent validation function before use
KPIs to instrument
- Hit rate (share of selected cases with a correction) against random selection
- No change rate of audits, before and after
- Additional tax assessed per audit hour
- Selection and hit rates across measurable groups
- Share of flags overridden by officers, with reasons
Human in the loop
Risk analysts own the models and the case lists; officers decide whether to open a check and carry out every compliance action. Taxpayers keep the normal review and appeal rights, and a fairness review of selection outcomes is reported to senior management at least yearly.
Common failure modes
- Discriminatory selection
- A model trained on past audits learns past bias, or uses a proxy such as nationality. The Dutch childcare benefits scandal and the IRS earned income tax credit disparities show the harm. Remove proxies and monitor outcomes by group.
- Feedback loops
- The model only learns from returns it selected, so it keeps finding the same risks. Random audits break the loop.
- Unexplainable flags
- Officers who cannot see why a return was flagged either ignore the model or trust it blindly. Show the drivers of each score.
- Blacklists that outlive their purpose
- Risk signals stored about individuals and never removed can follow people for years. Set retention and review rules for every signal list.
What are the risks and rules?
EU AI Act
Depends on design
Risk selection for administrative tax audits is not listed in Annex III, and Recital 59 says systems used by tax and customs authorities in administrative proceedings should not be treated as high risk law enforcement systems. Use in criminal tax investigations (Annex III point 6, law enforcement), or evaluating the eligibility of natural persons for public assistance benefits run through the tax system (Annex III point 5(a)), can make it high risk. When individuals are scored in administrative tax work, the GDPR applies, including its profiling rules (Member States may restrict some rights for taxation matters under Article 23). Article 22 applies when a decision with legal or similarly significant effect is taken solely by the model. Criminal investigations fall outside the GDPR and under the Law Enforcement Directive (EU) 2016/680 instead.
Rules that apply
Guidance
- Regulation (EU) 2024/1689 (AI Act), Recital 59 (European Union, Europe). Systems intended for administrative proceedings by tax and customs authorities should not be classified as high risk law enforcement systems.
- Algoritmeregister van de Nederlandse overheid (Government of the Netherlands, Europe). The Dutch national algorithm register, where the Belastingdienst and other agencies publish their selection and risk models with purpose, method and human oversight.
- Algorithmic Transparency Recording Standard hub (UK government, Europe). UK public bodies, including HMRC, publish transparency records for algorithmic tools used in compliance work.
Controls to put in place
- Register entry per model with purpose, data, owner and human oversight
- Fairness monitoring of selection and hit rates, reported at least yearly
- Explanation of each flag available to the officer and, on request, in disputes
- Independent model validation before first use and after material change
- Retention limits on risk signals held about individuals
When it went wrong elsewhere
- Amnesty International: Xenophobic machines: Discrimination through unregulated use of algorithms in the Dutch childcare benefits scandal. The Dutch tax authorities used a risk classification model in which nationality counted as a risk factor when checking childcare benefit applications, contributing to wrongful fraud accusations against parents.
- Stanford SIEPR: Measuring and mitigating racial disparities in tax audits. Researchers estimate that, despite race blind selection, Black taxpayers were audited at 2.9 to 4.7 times the rate of non Black taxpayers, driven mainly by audits of earned income tax credit claims.
Frequently asked questions
- Do tax administrations really use AI to choose audits?
- Yes, in varying forms. The IRS announced in 2023 that machine learning helped select large partnership returns for examination, HMRC gives around 5,500 VAT officers an anomaly detection tool, and the Dutch Belastingdienst publishes its selection models, such as a rules based VAT signal model for large businesses, in the national algorithm register.
- Is audit selection by AI high risk under the EU AI Act?
- Usually not for administrative tax audits, which Recital 59 keeps out of the law enforcement category. It can become high risk when used in criminal investigations or to decide on public benefits. When individuals are scored, GDPR profiling rules apply, with Article 22 covering any decision the model takes on its own.
- How do you keep selection fair?
- Exclude protected characteristics and proxies, keep a random sample, and compare selection and hit rates across groups every year. The childcare benefits scandal and research on US earned income tax credit audits show what happens without these checks.
How to cite this page
Blits.ai AI Use Case Library, "AI for tax compliance risk scoring and audit selection", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/tax-compliance-risk-scoring. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published