AI use case

Dynamic AML customer risk rating with machine learning

Explainable machine learning that produces the money laundering risk rating itself: it computes and continuously updates each customer's rating from due diligence data, products, geography, behaviour and screening results, and shows which factors drive the rating and when enhanced due diligence is warranted.

By Len Debets · Last verified 27 September 2026 · 1 public deployment

USD 120,000 to USD 1.3 million
Indicative value per year
A bank with 500,000 customers and 15,000 enhanced due diligence reviews a year. Worked example, see how it is calculated.

What problem does it solve?

Every bank must rate the money laundering risk of each customer and apply more scrutiny to the higher risk ones. In many banks the rating is a static scorecard filled in at onboarding and refreshed at fixed intervals; Fenergo describes periodic KYC checks as typically carried out annually or every two years. Between reviews the customer's behaviour can change completely without the rating moving.

Static scorecards can also drift out of line with reality: customers can end up high risk because of a single attribute, which adds work for enhanced due diligence teams, while risky behaviour in a low rated customer can go unnoticed until an alert or a law enforcement request. Supervisors expect a documented, risk based approach, and industry principles such as the Wolfsberg Group's ask that machine learning results can be explained from the data that went in.

How does it work?

  1. Combine the data. Due diligence attributes, products and channels, geographies, screening results, transaction behaviour and alert history are brought together per customer.
  2. Score with explanations. An interpretable model, or a model with feature attribution, produces a risk score and the factors that raise or lower it, alongside the bank's regulatory minimum rules (for example PEPs are always high risk).
  3. Update on events. The score is recalculated when behaviour or data changes, not only on the review date, and a material move creates a task.
  4. Route the work. Customers moving into higher risk bands are queued for enhanced due diligence with the drivers listed; moves down are reviewed before they reduce scrutiny.
  5. Govern the model. Rating distribution, stability and outcomes (alerts, reports, exits per band) are monitored, and the model is validated like any other risk model.
Audience
Back office
Autonomy
Supervised agent
Adoption
Early adopters
Channels
Internal tools, API and system to system

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

No public deployment has disclosed a measurable outcome yet.

Value drivers: Compliance quality, Risk and loss reduction, Employee productivity.

Indicative value

A bank with 500,000 customers and 15,000 enhanced due diligence reviews a year

USD 120,000 to USD 1.3 million

Enhanced due diligence effort redirected per year

How this is calculated

Formula: eddReviews * hoursPerReview * misratedShare * costPerHour. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Enhanced due diligence reviews per year eddReviews, reviews per year15,00015,000The reference bank.
Hours per enhanced due diligence review hoursPerReview, hours per review25Editorial assumption. Replace with your own.
Share of reviews avoided because customers were rated high only by static rules misratedShare, fraction of reviews0.10.25Editorial assumption. Replace with the results of your own rerating exercise.
Fully loaded analyst cost per hour costPerHour, USD per hour4070Editorial assumption. Replace with your own.

What it leaves out: Assumes the released effort is redirected to genuinely risky customers rather than cut. It leaves out the value of catching risk between reviews and the cost of model validation.

Who already uses it?

1 public deployment, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

bunq

Netherlands · Banking · 2022

ProductionGrade B

As described in a 2022 ruling, Dutch neobank bunq assigned new private customers a "regular user profile" derived from data analysis of its customer base, instead of asking each customer about the purpose of the account, and then monitored transaction behaviour to adjust that profile and raise the customer's risk scores when needed. bunq says it favours technology such as machine learning; the ruling mentions a machine learning model in transaction monitoring but does not show that machine learning produces the rating itself. In October 2022 the Dutch Trade and Industry Appeals Tribunal (CBb) ruled largely in bunq's favour in its dispute with De Nederlandsche Bank over this approach. In May 2025 De Nederlandsche Bank fined bunq EUR 2.6 million because it had not sufficiently followed up signals in four customer files it had itself identified as high risk; bunq has appealed, and the notice does not attribute the shortcomings to the data driven approach.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Customer due diligence attributes with data quality measures
  • Transaction behaviour aggregated per customer
  • Screening, alert and suspicious activity report history
  • The bank's enterprise wide money laundering risk assessment and rating methodology

Systems to integrate

  • KYC and customer lifecycle management system
  • Transaction monitoring and case management
  • Screening engines
  • Core banking and product systems
  • Periodic review and enhanced due diligence workflow

Complexity: High

The rating drives regulatory obligations for every customer, so it must be explainable, stable, validated and consistent with the bank's documented risk assessment. Data joining across KYC, transactions and screening is usually the biggest piece of work.

  1. 1

    Keep the regulatory floor explicit

    Write down the ratings that rules dictate (for example PEPs and high risk jurisdictions) and keep them as hard constraints above the model.

  2. 2

    Choose explainability over marginal accuracy

    Prefer a model whose drivers an analyst can read out to an examiner. A slightly better score that nobody can explain is not usable.

  3. 3

    Back test against outcomes

    Check that higher bands contain more alerts, reports and exits than lower ones, on history, and compare with the current scorecard.

  4. 4

    Rerate in parallel

    Rerate the whole book in parallel and review the largest moves with compliance before switching, including customers that would move down.

  5. 5

    Turn on event driven updates

    Recalculate on material events and route moves into higher bands to enhanced due diligence with the drivers attached.

Guardrails

  • Regulatory minimum ratings enforced as hard rules above the model
  • Every rating shows its contributing factors in plain language
  • Moves to a lower band that reduce scrutiny are reviewed by a human
  • Bias testing so that nationality or other protected characteristics do not drive ratings beyond what the risk assessment justifies
  • Model inventory entry with validation and stability monitoring

KPIs to instrument

  • Distribution of customers across bands and month on month stability
  • Alerts, reports and exits per band (the rating should separate them)
  • Enhanced due diligence volume and time per review
  • Share of rating moves overturned by analysts
  • Time from a material event to the updated rating

Human in the loop

The model rates and updates; analysts confirm moves into and out of high risk, and compliance owns the methodology, approves the model and reviews distribution and outcome reports. Decisions to restrict or exit a customer remain human decisions.

Common failure modes

The black box rating
A rating the bank cannot explain is hard to defend to a supervisor and gives analysts nothing to investigate, however accurate it is. Use interpretable models or reliable attribution.
Rating churn
Customers flip between bands with every transaction, creating work without insight. Smooth scores and use materiality thresholds.
Proxy discrimination
The model leans on nationality or postcode beyond what the risk assessment supports. Test and constrain features.

What are the risks and rules?

EU AI Act

Depends on design

An AML customer risk rating is not listed in Annex III. Article 5(1)(d) prohibits AI risk assessments that predict whether a natural person will commit or will likely commit a criminal offence based solely on profiling of that person or on assessing their personality traits and characteristics; it exempts only AI that supports the human assessment of a person's involvement in a criminal activity, which is already based on objective and verifiable facts directly linked to a criminal activity. An AML customer risk rating built from due diligence attributes, transaction behaviour and screening results is itself an automated evaluation of a person's situation and behaviour, which is profiling under GDPR Article 4(4), and due diligence facts such as occupation, geography and products are not facts directly linked to a criminal activity, so the rating does not sit squarely inside the exemption. What keeps it a defensible AML due diligence tool rather than an offence prediction is that it does not itself accuse a person of an offence: it sets a level of scrutiny, a human analyst reviews material moves, and regulatory minimum rules sit above the model as hard constraints. A rating driven mainly by nationality or other personal attributes weakens that position further, which is why the proxy discrimination guardrail matters. If the same score is used to evaluate the creditworthiness of natural persons or to establish their credit score, that use falls under Annex III point 5(b) and is high risk, so keep the AML rating and credit decisions separate.

Guidance

Controls to put in place

  • Documented rating methodology linked to the enterprise wide risk assessment
  • Factor level explanation stored with every rating
  • Independent validation, stability monitoring and annual review of the model
  • Bias testing on protected characteristics
  • Human review of material moves in both directions

When it went wrong elsewhere

  • De Nederlandsche Bank fines bunq for insufficient customer due diligence. On 6 May 2025 the Dutch central bank fined bunq EUR 2.6 million because it did not sufficiently follow up signals and transaction monitoring alerts in four customer files it had itself identified as high risk (period January 2021 to May 2022); bunq has appealed. The notice does not blame bunq's data driven approach, but it shows that a rating is only as good as the human follow up it triggers.

Frequently asked questions

What is a dynamic customer risk rating?
A money laundering risk rating that updates when the customer's data or behaviour changes, instead of only at a scheduled review, and shows which factors drive it.
Does a machine learning risk rating have to be explainable?
In practice yes. The Wolfsberg Group's principles for AI and machine learning in financial crime compliance ask that results "can be adequately explained or proven given the data inputs", and analysts need the drivers to run enhanced due diligence. Choose interpretable models or reliable attribution over a small gain in accuracy.
Is the risk rating high risk under the EU AI Act?
Not as an AML tool: AML risk rating is not listed in Annex III. It becomes high risk if the same score is used to evaluate the creditworthiness of natural persons or to set their credit score (Annex III point 5(b)), so keep the uses separate and documented. Article 5(1)(d) bans predicting that a person will commit an offence from profiling or personality traits alone; it is designed as a due diligence tool, not an offence prediction, so do not let it drive a rating from nationality or other personal traits alone, keep the regulatory minimum rules and analyst review in place, and treat the rating itself as a level of scrutiny rather than an accusation.
How is this different from a machine learning transaction monitoring score?
A monitoring score such as HSBC's Dynamic Risk Assessment, which Google Cloud describes as a customer risk score offered as an alternative to rules based transaction alerting, decides which customers investigators look at for suspicious activity (see AML alert triage). The rating on this page decides the level of due diligence each customer gets. The two share data and methods, so HSBC's deployment is related evidence here rather than a direct example.

How to cite this page

Blits.ai AI Use Case Library, "Dynamic AML customer risk rating with machine learning", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/dynamic-customer-risk-rating. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

BankingPayments and cards

AI for perpetual KYC and event driven customer due diligence

AI that keeps each customer's due diligence file current by replacing calendar driven KYC reviews with continuous, event driven refreshes: it watches for trigger events such as a change of ownership, address, behaviour or a new adverse finding, refreshes the file automatically where it can, and involves an analyst only when something material has changed. The risk rating itself and the first file for a new business client are separate use cases.

Deployments
5 public, best grade B
Reported cost reduction
40%
JPMorgan Chase, organization claim
BankingPayments and cards

AI for AML transaction monitoring alert triage

Machine learning and AI agents that score anti money laundering alerts for genuine risk, close clear false positives with a written and stored rationale, and hand investigators the remaining alerts already enriched with the customer, counterparty and transaction context.

Deployments
8 public, best grade B
Reported false positive reduction
86%
Shift4, vendor claim
BankingPayments and cards

AI for PEP and adverse media screening

AI that continuously scans news, court records, registries and other open sources in many languages for negative information and political exposure linked to customers, counterparties and beneficial owners, discards look alikes, and summarises credible risk for the analyst with the sources attached.

Deployments
7 public, best grade B
Reported handling time reduction
at least 60%
Save the Children, vendor claim
BankingPayments and cards

AI for business onboarding (KYB) and beneficial ownership discovery

An AI agent that builds the know your business (KYB) due diligence file for a new or reviewed corporate client, before any account is opened: it collects registry, incorporation and ownership documents, resolves the entity across sources, maps the ownership chain through holding companies, nominees and trusts to the ultimate beneficial owners, screens the entity and its owners, and presents a risk scored case for a compliance analyst to decide.

Deployments
3 public, best grade C
Reported automation rate
25%
BNY, organization claim
Wealth and asset managementBanking

AI agent for source of wealth due diligence in private banking

An AI agent that reads a prospective private client's documents, extracts and corroborates how their wealth was built, checks plausibility against benchmarks and external sources, and drafts the source of wealth and enhanced due diligence narrative for the relationship manager and compliance analyst, who decide on the risk rating and the relationship.

Deployments
3 public, best grade B
Autonomy
Copilot