AI use case

AI for benefit fraud and error detection in social security

Risk models that help a social security or benefits agency decide which claims, payments and recipients to check for fraud or error, so that caseworkers verify the riskiest cases first, while every decision on entitlement stays with a person and the model is tested for fairness before and during use.

By Len Debets · Last verified 27 September 2026 · 5 public deployments

2.5x
Reported detection improvement
Department for Work and Pensions, organization claim.
EUR 500,000 to EUR 13.5 million
Indicative value per year
A national benefits agency paying EUR 2 billion a year in a benefit with known fraud and error risk. Worked example, see how it is calculated.

What problem does it solve?

Benefits agencies pay very large sums to millions of people, and some of it goes to the wrong place: organised fraud, individual misrepresentation, and honest mistakes by claimants or by the agency itself. Checking everyone is impossible and would delay payments to people who need them urgently, so agencies have to choose whom to check.

That choice is where automated risk scoring has done serious harm. The Dutch childcare benefits scandal, the SyRI welfare fraud system that a Dutch court stopped in 2020, Rotterdam's welfare risk model and Australia's Robodebt scheme each show one or more of the same failures: selection on nationality or on proxies for it, systems too opaque to check or challenge, and a burden of proof shifted onto people who were then treated as if they owed money or had committed fraud. Under the EU AI Act, systems that evaluate eligibility for public assistance, or grant, reduce, revoke or reclaim it, are high risk. The job is not only to catch fraud but to do so lawfully, proportionately and transparently.

How does it work?

  1. Define the risk precisely. A model targets one defined risk at one point in the process, for example an advance request before payment or a specific change in circumstances, not "fraud" in general.
  2. Score at the point of decision. The claim data is scored in real time or in batch. The output is a referral for a check, never a decision on entitlement.
  3. Blind and mix the referrals. High risk referrals go to caseworkers together with a random control group, and the caseworker is not told which is which, so the check is not biased by the score.
  4. Verify with the person. A caseworker reviews the evidence and, where needed, asks the claimant, then decides. Declines follow the normal notice, review and appeal route.
  5. Measure effectiveness and fairness. Confirmed fraud and error rates in model referrals are compared with the random group, overall and by group, and the model is retrained or stopped when it targets groups without a matching rate of confirmed findings.
Audience
Back office
Autonomy
Assist
Adoption
Early adopters
Channels
API and system to system, Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI for benefit fraud and error detection in social security
KPIMedianReported rangeData pointsClaimed by
Detection improvementToo few to pool
2.5x
11 organization

Value drivers: Risk and loss reduction, Compliance quality, Lower cost to serve.

Indicative value

A national benefits agency paying EUR 2 billion a year in a benefit with known fraud and error risk

EUR 500,000 to EUR 13.5 million

Fraud and error losses prevented per year

How this is calculated

Formula: paymentsValue * lossRate * preventedShare. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Value of payments in scope paymentsValue, EUR per year1,000,000,0003,000,000,000Editorial assumption. Replace with the payments the model will actually screen.
Share of payments lost to fraud and error lossRate, fraction of payments0.010.03Editorial assumption. Use your own official fraud and error statistics, which vary widely by benefit.
Share of those losses prevented by better targeted checks preventedShare, fraction of losses0.050.15Editorial assumption, deliberately modest. DWP reports its advances model is 2.5 times more effective than random selection, but a better hit rate prevents only part of the loss.

What it leaves out: Gross losses prevented only. It leaves out the cost of the checks, the cost of wrongly delayed or refused payments to legitimate claimants, legal and reputational risk, and the cost of the governance a high risk system requires.

Who already uses it?

5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Department for Work and Pensions

United Kingdom · Government and public sector · 2026

ScaledGrade B

DWP scores requests for Universal Credit advances in real time with a supervised machine learning classifier and refers the highest risk requests to a caseworker before payment. The caseworker is not told the referral came from the model, a random control group is referred alongside, and every decision to decline is made by a person and can be appealed. DWP's published effectiveness assessment for April 2025 to March 2026 finds the model 2.5 times more effective than random selection with a median payment delay of one day for approved referrals. It also finds that non UK nationals and several age bands were referred more often without a matching increase in confirmed fraud, and that referrals of couples were less often confirmed than those of single claimants; a retrained model was being tested.

  • Detection improvement: 2.5x, 1 April 2025 to 31 March 2026
    "The performance information for 2025 to 2026 demonstrates the model is 2.5 times more effective at identifying fraud risk than a randomised control group sample."
    Claimed by: organization

U.S. Department of the Treasury, Bureau of the Fiscal Service

United States · Government and public sector · 2024

ScaledGrade B

The Office of Payment Integrity in Treasury's Bureau of the Fiscal Service added machine learning and risk based screening to how it checks federal payments. Treasury reports that these enhanced processes prevented and recovered over USD 4 billion in fraud and improper payments in fiscal year 2024, up from USD 652.7 million the year before. The release names machine learning only for identifying Treasury check fraud, which led to USD 1 billion in recovery; the USD 2.5 billion it reports from identifying and prioritising high risk transactions is not described as machine learning.

  • Fraud losses prevented: USD 1 billion, fiscal year 2024 (October 2023 to September 2024)
    "Expediting the identification of Treasury check fraud with machine learning AI resulting in $1 billion in recovery."
    Claimed by: organization

Gemeente Rotterdam

Netherlands · Government and public sector · 2022

Paused or rolled backGrade B

From 2017 the City of Rotterdam used a machine learning model that gave each social assistance recipient a risk score between 0 and 1 for receiving benefits they were not, or no longer, entitled to, based on the outcomes of earlier eligibility reviews. High scores were one route to an invitation for a review interview, and an income consultant decided the outcome. The city classified the model as high risk and stopped using it in early 2022; its register entry states that a review found it was not currently possible to build a risk model that fits the city's policy. The register says the model processed no nationality, age or health data; journalists who obtained the model file reported 315 inputs, including age, gender and language skills, and found that it discriminated by ethnicity, age, gender and parenthood.

No outcome disclosed.

Uitvoeringsinstituut Werknemersverzekeringen (UWV)

Netherlands · Government and public sector · 2022

ProductionGrade B

The Dutch employee insurance agency UWV uses a risk scan on applications for unemployment benefit (WW) to signal cases where the applicant may have become unemployed through their own fault, which would remove the entitlement. The scan combines several data points, never a single characteristic, and uses no personal characteristics such as origin, gender or age; high risk applications are offered to staff for a fuller investigation, and staff decide. Specialists check the input data quality monthly and whether the development population is still representative, and an independent party reviews the work when the scan is developed further. Randomly chosen applications (30 percent, according to the register) are added to the scan's selection, so staff do not know which cases the scan flagged. UWV carried out a DPIA and an ethical impact assessment. In use since August 2022; no outcome figures are published.

No outcome disclosed.

Centers for Medicare & Medicaid Services

United States · Government and public sector · 2021

AnnouncedGrade B

The CMS Division of Payment Reconciliation monitors Prescription Drug Event (PDE) data from Medicare Part D for accuracy. It reported an AI project to identify outliers in that data so that errors can be corrected through outreach to the plans and improper payments, an overpayment or an underpayment of a government benefit, are prevented. The project was at the initiated stage in the 2024 federal inventory; no results are published.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Confirmed outcomes of past checks, including checks that found nothing
  • Claim data available at the decision point, with a legal basis for each item
  • Protected characteristic data or reliable proxies for fairness testing

Systems to integrate

  • Benefit claim and payment systems
  • Caseworker case management
  • Notice, review and appeal processes
  • Data warehouse for monitoring and evaluation

Complexity: High

Technically this is a scoring model; in practice it is one of the most regulated and contested uses of AI in government. It needs a clear legal basis, a data protection impact assessment, an equality or fundamental rights assessment, published documentation and an evaluation design with a random control group.

  1. 1

    Choose a narrow, well evidenced risk

    Start with one payment type where fraud is documented, as DWP did with Universal Credit advances, rather than a general score of every recipient.

  2. 2

    Do the rights assessment first

    Complete the data protection impact assessment and an equality or fundamental rights impact assessment before build, and publish a transparency record.

  3. 3

    Build the evaluation into the design

    Keep a random control group of referrals and hide the source of each referral from caseworkers, so effectiveness and fairness can be measured honestly.

  4. 4

    Test disparities in referral and in outcome

    For every group you can measure, compare how often the model refers people with how often those referrals are confirmed. DWP's assessment found groups referred more often without more confirmed fraud, and DWP retrained the model, which was still in testing when it reported.

  5. 5

    Define the stop rule

    Agree in advance what finding would pause the model. Rotterdam stopped its welfare model when it concluded it could not currently build a model that fits its policy.

Guardrails

  • The model only refers cases for a check; it never decides, reduces or stops a payment
  • No protected characteristic, nationality or obvious proxy as a feature
  • Random control group and blinded referrals for every model in use
  • Published transparency record and effectiveness assessment at least yearly
  • Payment timeliness for legitimate claimants monitored as a harm measure

KPIs to instrument

  • Confirmed fraud and error rate in model referrals versus the random control group
  • Referral and outcome disparities by group
  • Payment delay for legitimate claimants who were referred
  • Share of declines overturned on review or appeal
  • Losses prevented per check

Human in the loop

Caseworkers review every referral and make every entitlement decision, with the claimant able to explain their circumstances. Claimants keep the normal review and appeal rights. A senior owner signs off the yearly effectiveness and fairness assessment and can suspend the model.

Common failure modes

Discrimination through proxies
Nationality is used directly, or language, neighbourhood or family status stand in for protected characteristics. The Dutch childcare benefits model used nationality as a risk factor, and journalists found Rotterdam's model scored on age, gender and language skills.
Automation reversing the burden of proof
When a score or data match is treated as proof, people must disprove a debt. Robodebt raised debts from averaged income data without other evidence and put the onus on recipients to contradict them. A person must establish the facts.
Secrecy that blocks accountability
Agencies that refuse to disclose how their models work cannot be challenged or corrected, as investigations in Sweden and Denmark have shown. Publish the documentation.
Chasing small sums at high human cost
Aggressive thresholds generate many checks on honest claimants for little recovered money. Measure the harm side alongside the savings.

What are the risks and rules?

EU AI Act

High risk

Annex III point 5(a): AI systems used by or on behalf of public authorities to evaluate the eligibility of natural persons for essential public assistance benefits and services, or to grant, reduce, revoke or reclaim them. A fundamental rights impact assessment (Article 27) is required before a public body deploys it. A design that scores people over time on their social behaviour or personal characteristics and leads to unrelated or disproportionate detrimental treatment would fall under the Article 5(1)(c) prohibition on social scoring.

Guidance

  • Algorithmic Transparency Recording Standard hub (UK government, Europe). UK central government bodies publish records of their algorithmic tools here, including DWP's record for its Universal Credit advances model.
  • Algoritmeregister van de Nederlandse overheid (Government of the Netherlands, Europe). Dutch government bodies register algorithms used in benefit control here, for example UWV's unemployment benefit risk scan and Rotterdam's welfare risk model, which is listed as no longer in use.

Controls to put in place

  • Fundamental rights and data protection impact assessments before deployment
  • Yearly published effectiveness and fairness assessment with a random control group
  • Blinded referrals and human decision on every case
  • Documented stop rule and senior owner with authority to suspend
  • Notice, review and appeal routes that do not depend on knowing the model exists

When it went wrong elsewhere

  • Amnesty International: Xenophobic machines, the Dutch childcare benefits scandal. The Dutch tax authorities used an algorithmic system to create risk profiles of childcare benefit applicants, with nationality as one of its risk factors. Amnesty found this resulted in discrimination and racial profiling.
  • District Court of The Hague: SyRI judgment (ECLI:NL:RBDHA:2020:865). The court ruled in February 2020 that the legislation for SyRI, a Dutch government system that linked data to flag welfare fraud risk, breached the right to private life because it was insufficiently transparent and verifiable.
  • Royal Commission into the Robodebt Scheme: report volume 1 (archived). Australia's Robodebt scheme raised welfare debts largely through income averaging, without other evidence, and placed the onus on recipients to contradict the result. The Royal Commission found the method neither produced accurate results nor complied with the income calculation provisions of the Social Security Act 1991. In 2020 the government decided to refund debts raised wholly or partly through averaging that had been repaid, and to reduce unpaid ones to zero, and in November 2020 it settled a class action. The report was presented on 7 July 2023. It was automation rather than machine learning, but the lesson on burden of proof applies.
  • Lighthouse Reports: Suspicion Machines (Rotterdam). Reconstruction of Rotterdam's welfare fraud model, which took 315 inputs including age, gender and language skills; the investigation found it discriminated by ethnicity, age, gender and parenthood.
  • Lighthouse Reports: Sweden's Suspicion Machine. Analysis of the Swedish Social Insurance Agency's fraud prediction algorithm for temporary child support found it disproportionately flagged women, migrants, low income earners and people without a university education.
  • Amnesty International: Denmark's AI powered welfare system fuels mass surveillance. Amnesty's 2024 report on the fraud control algorithms of Denmark's welfare authority, Udbetaling Danmark, warns of mass surveillance and discrimination against marginalised groups, including through inputs on "foreign affiliation". The authority refused full access to the code and data.

Frequently asked questions

Is benefit fraud detection high risk under the EU AI Act?
Yes, when the system evaluates eligibility for public assistance or is used to grant, reduce, revoke or reclaim benefits (Annex III point 5(a)). Public bodies must also carry out a fundamental rights impact assessment before use.
Does it work?
It can improve targeting. DWP reports its Universal Credit advances model is 2.5 times more effective than random selection in 2025 to 2026, with a median payment delay of one day for approved referrals. The same assessment found disparities by nationality, age and couple status, which is why the random control group and yearly assessment matter.
What went wrong in the Dutch, Australian and Rotterdam cases?
Selection used nationality or proxies such as language skills, people had to disprove a score or data match, and systems were hard to scrutinise. Rotterdam stopped its model in 2022, a Dutch court struck down the SyRI legislation in 2020, and debts Robodebt raised through income averaging were refunded or reduced to zero before a Royal Commission reported in 2023.

How to cite this page

Blits.ai AI Use Case Library, "AI for benefit fraud and error detection in social security", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/benefit-fraud-and-error-detection. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Government and public sector

AI for tax compliance risk scoring and audit selection

Models that score tax returns, taxpayers and transactions for the risk of error, underreporting or fraud, so that a tax administration spends its audit and compliance capacity where the risk is highest, with an officer deciding every compliance action and the selection itself monitored for fairness.

Deployments
3 public, best grade B
Autonomy
Assist
Government and public sector

AI assistant for benefits eligibility questions and applications

An AI assistant that helps people understand which public benefits and grants may apply to them, explains the rules and documents in plain language, guides them through the application and checks it for completeness, while the eligibility decision stays with the agency's rules and caseworkers.

Deployments
7 public, best grade B
Reported accuracy
97%
Department for Work and Pensions, organization claim
BankingPayments and cards

AI for application and identity fraud detection

AI that checks incoming account and loan applications for forged or AI generated documents, synthetic and stolen identities, and coordinated application rings, by analysing documents, device and application data across the whole queue and cross checking against bureau and official sources.

Deployments
6 public, best grade B
Reported detection improvement
2.5x
Department for Work and Pensions, organization claim
Cross industryBanking

AI system and model inventory with shadow AI discovery

A governed register of every AI system and model an organization builds, buys or uses, with its owner, purpose, data, risk tier and approval status, kept current by AI that discovers unregistered use, reads the documentation and assembles the evidence a board, auditor or supervisor asks for.

Deployments
4 public, best grade B
Autonomy
Copilot