What problem does it solve?
Every bank must rate the money laundering risk of each customer and apply more scrutiny to the higher risk ones. In many banks the rating is a static scorecard filled in at onboarding and refreshed at fixed intervals; Fenergo describes periodic KYC checks as typically carried out annually or every two years. Between reviews the customer's behaviour can change completely without the rating moving.
Static scorecards can also drift out of line with reality: customers can end up high risk because of a single attribute, which adds work for enhanced due diligence teams, while risky behaviour in a low rated customer can go unnoticed until an alert or a law enforcement request. Supervisors expect a documented, risk based approach, and industry principles such as the Wolfsberg Group's ask that machine learning results can be explained from the data that went in.
- A Fenergo study found that more than half of financial institutions spend between 61 and 150 days on client KYC reviews, at an average cost of $2,200 per review.Ongoing Customer Due Diligence with Perpetual KYC (2026)
How does it work?
- Combine the data. Due diligence attributes, products and channels, geographies, screening results, transaction behaviour and alert history are brought together per customer.
- Score with explanations. An interpretable model, or a model with feature attribution, produces a risk score and the factors that raise or lower it, alongside the bank's regulatory minimum rules (for example PEPs are always high risk).
- Update on events. The score is recalculated when behaviour or data changes, not only on the review date, and a material move creates a task.
- Route the work. Customers moving into higher risk bands are queued for enhanced due diligence with the drivers listed; moves down are reviewed before they reduce scrutiny.
- Govern the model. Rating distribution, stability and outcomes (alerts, reports, exits per band) are monitored, and the model is validated like any other risk model.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Early adopters
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
No public deployment has disclosed a measurable outcome yet.
Value drivers: Compliance quality, Risk and loss reduction, Employee productivity.
Indicative value
A bank with 500,000 customers and 15,000 enhanced due diligence reviews a year
USD 120,000 to USD 1.3 million
Enhanced due diligence effort redirected per year
How this is calculated
Formula: eddReviews * hoursPerReview * misratedShare * costPerHour. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Enhanced due diligence reviews per year eddReviews, reviews per year | 15,000 | 15,000 | The reference bank. |
| Hours per enhanced due diligence review hoursPerReview, hours per review | 2 | 5 | Editorial assumption. Replace with your own. |
| Share of reviews avoided because customers were rated high only by static rules misratedShare, fraction of reviews | 0.1 | 0.25 | Editorial assumption. Replace with the results of your own rerating exercise. |
| Fully loaded analyst cost per hour costPerHour, USD per hour | 40 | 70 | Editorial assumption. Replace with your own. |
What it leaves out: Assumes the released effort is redirected to genuinely risky customers rather than cut. It leaves out the value of catching risk between reviews and the cost of model validation.
Who already uses it?
1 public deployment, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
bunq
Netherlands · Banking · 2022
As described in a 2022 ruling, Dutch neobank bunq assigned new private customers a "regular user profile" derived from data analysis of its customer base, instead of asking each customer about the purpose of the account, and then monitored transaction behaviour to adjust that profile and raise the customer's risk scores when needed. bunq says it favours technology such as machine learning; the ruling mentions a machine learning model in transaction monitoring but does not show that machine learning produces the rating itself. In October 2022 the Dutch Trade and Industry Appeals Tribunal (CBb) ruled largely in bunq's favour in its dispute with De Nederlandsche Bank over this approach. In May 2025 De Nederlandsche Bank fined bunq EUR 2.6 million because it had not sufficiently followed up signals in four customer files it had itself identified as high risk; bunq has appealed, and the notice does not attribute the shortcomings to the data driven approach.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Customer due diligence attributes with data quality measures
- Transaction behaviour aggregated per customer
- Screening, alert and suspicious activity report history
- The bank's enterprise wide money laundering risk assessment and rating methodology
Systems to integrate
- KYC and customer lifecycle management system
- Transaction monitoring and case management
- Screening engines
- Core banking and product systems
- Periodic review and enhanced due diligence workflow
Complexity: High
The rating drives regulatory obligations for every customer, so it must be explainable, stable, validated and consistent with the bank's documented risk assessment. Data joining across KYC, transactions and screening is usually the biggest piece of work.
- 1
Keep the regulatory floor explicit
Write down the ratings that rules dictate (for example PEPs and high risk jurisdictions) and keep them as hard constraints above the model.
- 2
Choose explainability over marginal accuracy
Prefer a model whose drivers an analyst can read out to an examiner. A slightly better score that nobody can explain is not usable.
- 3
Back test against outcomes
Check that higher bands contain more alerts, reports and exits than lower ones, on history, and compare with the current scorecard.
- 4
Rerate in parallel
Rerate the whole book in parallel and review the largest moves with compliance before switching, including customers that would move down.
- 5
Turn on event driven updates
Recalculate on material events and route moves into higher bands to enhanced due diligence with the drivers attached.
Guardrails
- Regulatory minimum ratings enforced as hard rules above the model
- Every rating shows its contributing factors in plain language
- Moves to a lower band that reduce scrutiny are reviewed by a human
- Bias testing so that nationality or other protected characteristics do not drive ratings beyond what the risk assessment justifies
- Model inventory entry with validation and stability monitoring
KPIs to instrument
- Distribution of customers across bands and month on month stability
- Alerts, reports and exits per band (the rating should separate them)
- Enhanced due diligence volume and time per review
- Share of rating moves overturned by analysts
- Time from a material event to the updated rating
Human in the loop
The model rates and updates; analysts confirm moves into and out of high risk, and compliance owns the methodology, approves the model and reviews distribution and outcome reports. Decisions to restrict or exit a customer remain human decisions.
Common failure modes
- The black box rating
- A rating the bank cannot explain is hard to defend to a supervisor and gives analysts nothing to investigate, however accurate it is. Use interpretable models or reliable attribution.
- Rating churn
- Customers flip between bands with every transaction, creating work without insight. Smooth scores and use materiality thresholds.
- Proxy discrimination
- The model leans on nationality or postcode beyond what the risk assessment supports. Test and constrain features.
What are the risks and rules?
EU AI Act
Depends on design
An AML customer risk rating is not listed in Annex III. Article 5(1)(d) prohibits AI risk assessments that predict whether a natural person will commit or will likely commit a criminal offence based solely on profiling of that person or on assessing their personality traits and characteristics; it exempts only AI that supports the human assessment of a person's involvement in a criminal activity, which is already based on objective and verifiable facts directly linked to a criminal activity. An AML customer risk rating built from due diligence attributes, transaction behaviour and screening results is itself an automated evaluation of a person's situation and behaviour, which is profiling under GDPR Article 4(4), and due diligence facts such as occupation, geography and products are not facts directly linked to a criminal activity, so the rating does not sit squarely inside the exemption. What keeps it a defensible AML due diligence tool rather than an offence prediction is that it does not itself accuse a person of an offence: it sets a level of scrutiny, a human analyst reviews material moves, and regulatory minimum rules sit above the model as hard constraints. A rating driven mainly by nationality or other personal attributes weakens that position further, which is why the proxy discrimination guardrail matters. If the same score is used to evaluate the creditworthiness of natural persons or to establish their credit score, that use falls under Annex III point 5(b) and is high risk, so keep the AML rating and credit decisions separate.
Rules that apply
Guidance
- Supporting Artificial Intelligence Adoption in AML/CFT (Hong Kong Monetary Authority, Asia Pacific). Describes banks moving from rules to holistic, data driven AML approaches and the supervisor's support programme.
- Principles for Using Artificial Intelligence and Machine Learning in Financial Crime Compliance (Wolfsberg Group, Global). Industry principles on accountability, openness and transparency that apply directly to a risk rating model.
- Notice 626 Prevention of Money Laundering and Countering the Financing of Terrorism, Banks (Monetary Authority of Singapore, Asia Pacific). Example of national rules on risk based customer due diligence and enhanced measures for higher risk customers.
Controls to put in place
- Documented rating methodology linked to the enterprise wide risk assessment
- Factor level explanation stored with every rating
- Independent validation, stability monitoring and annual review of the model
- Bias testing on protected characteristics
- Human review of material moves in both directions
When it went wrong elsewhere
- De Nederlandsche Bank fines bunq for insufficient customer due diligence. On 6 May 2025 the Dutch central bank fined bunq EUR 2.6 million because it did not sufficiently follow up signals and transaction monitoring alerts in four customer files it had itself identified as high risk (period January 2021 to May 2022); bunq has appealed. The notice does not blame bunq's data driven approach, but it shows that a rating is only as good as the human follow up it triggers.
Frequently asked questions
- What is a dynamic customer risk rating?
- A money laundering risk rating that updates when the customer's data or behaviour changes, instead of only at a scheduled review, and shows which factors drive it.
- Does a machine learning risk rating have to be explainable?
- In practice yes. The Wolfsberg Group's principles for AI and machine learning in financial crime compliance ask that results "can be adequately explained or proven given the data inputs", and analysts need the drivers to run enhanced due diligence. Choose interpretable models or reliable attribution over a small gain in accuracy.
- Is the risk rating high risk under the EU AI Act?
- Not as an AML tool: AML risk rating is not listed in Annex III. It becomes high risk if the same score is used to evaluate the creditworthiness of natural persons or to set their credit score (Annex III point 5(b)), so keep the uses separate and documented. Article 5(1)(d) bans predicting that a person will commit an offence from profiling or personality traits alone; it is designed as a due diligence tool, not an offence prediction, so do not let it drive a rating from nationality or other personal traits alone, keep the regulatory minimum rules and analyst review in place, and treat the rating itself as a level of scrutiny rather than an accusation.
- How is this different from a machine learning transaction monitoring score?
- A monitoring score such as HSBC's Dynamic Risk Assessment, which Google Cloud describes as a customer risk score offered as an alternative to rules based transaction alerting, decides which customers investigators look at for suspicious activity (see AML alert triage). The rating on this page decides the level of due diligence each customer gets. The two share data and methods, so HSBC's deployment is related evidence here rather than a direct example.
How to cite this page
Blits.ai AI Use Case Library, "Dynamic AML customer risk rating with machine learning", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/dynamic-customer-risk-rating. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published