What problem does it solve?
Every complaint is handled one by one, but the reason it happened is rarely unique. The same unclear letter, broken app journey or misapplied fee produces hundreds of complaints, spread across phone notes, emails, chat logs and ombudsman referrals, each coded slightly differently by a different handler. Complaint categories are built for handling and reporting volumes, not for finding causes, so management information shows how many complaints arrived, not why.
Regulators expect more. UK rules require firms to identify and remedy recurring or systemic problems, and the FCA's review of complaints and root cause analysis at 40 firms found that firms did not always measure whether their fixes worked, and that some complaints reports appeared to be prepared for operational purposes such as resourcing without also looking at how to improve customer outcomes. Manual root cause work at this volume often relies on samples, and the theme that matters most can be the one nobody sampled.
- The FCA's thematic review of complaints and root cause analysis at 40 firms found that firms did not always measure the impact of the changes they made after finding a root cause.Complaints and root cause analysis: good practice and areas for improvement (2024)
How does it work?
- Collect every complaint. Complaint records, call and chat transcripts, emails and ombudsman cases land in one store with product, channel, outcome and customer segment.
- Normalise and deduplicate. The AI summarises each complaint in a standard form, removes duplicates about the same event and tags vulnerability signals.
- Cluster into themes. Complaints are grouped by what actually went wrong, not by the code a handler picked, and each theme gets a plain language description with example cases.
- Trace to a cause. Retrieval over process maps, product terms, change logs and the control library links each theme to the likely process, product change or control failure, and flags themes that grow, spread across products or hit vulnerable customers.
- Validate and assign. A root cause analyst validates the theme and the cause on a sample, then assigns an owner in the business, not the complaints team.
- Track the fix. Actions, owners and dates are tracked, and the AI measures whether the theme shrinks after the fix, which is the evidence regulators ask for.
- Audience
- Back office
- Autonomy
- Copilot
- Adoption
- Emerging
- Channels
- Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
No public deployment has disclosed a measurable outcome yet.
Value drivers: Compliance quality, Customer experience, Risk and loss reduction, Lower cost to serve.
Indicative value
A retail bank receiving 100,000 complaints a year
USD 300,000 to USD 2.4 million
Complaint handling cost avoided by fixing systemic causes per year
How this is calculated
Formula: complaints * systemicShare * avoided * costPerComplaint. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Complaints received per year complaints, complaints per year | 100,000 | 100,000 | The reference bank. Replace with your own volume. |
| Share of complaints linked to a fixable systemic cause systemicShare, fraction of complaints | 0.1 | 0.2 | Editorial assumption. Replace with the share your own root cause work attributes to recurring causes. |
| Share of those complaints avoided after the cause is fixed avoided, fraction of systemic complaints | 0.2 | 0.4 | Editorial assumption, replace with your own. No public benchmark exists yet. |
| Fully loaded cost of handling one complaint costPerComplaint, USD per complaint | 150 | 300 | Editorial assumption covering handling time and review, excluding redress. Replace with your own. |
What it leaves out: Handling cost only. It leaves out redress and remediation programmes avoided, ombudsman fees, the customer and conduct benefit and the cost of the fixes themselves, and it assumes the business acts on the insight.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Centers for Medicare and Medicaid Services
United States · Government and public sector · 2024
A team at the US Centers for Medicare and Medicaid Services is piloting AI that analyses a high volume of complaint cases to identify root causes and trends, maps them to the applicable regulatory citations, draws a sample for subject matter experts to validate and recommends next steps. The stated aim is to reduce repeat issues that delay benefits or access to care and to improve health plan compliance. The 2024 inventory describes an earlier proof of concept that used a large language model on complaint data used by the same office (OPOLE). No results are published.
No outcome disclosed.
Board of Governors of the Federal Reserve System
United States · Government and public sector · 2019
The Federal Reserve Board's Division of Consumer and Community Affairs has used an in house natural language processing tool since 2019 to sort large volumes of consumer complaint narratives into topics, so staff can analyse and respond to them. For each narrative it outputs a topic number, a fit score and the top five terms of that topic. The input is complaint data from the CFPB. It is a central bank analysing consumer complaints about financial companies from the CFPB database rather than a firm analysing its own complaints, but the method is the same clustering step a bank's root cause work starts from.
No outcome disclosed.
Federal Trade Commission
United States · Government and public sector · 2019
Since 2019 the US Federal Trade Commission has used AI on the complaints it receives through ReportFraud and other channels: one model classifies uncategorised complaints by product and service code, another groups duplicate complaints about the same issue so investigators can see which entities attract multiple reports, alongside graph analytics that connect complaints about the same company when it uses different names, phone numbers or aliases. It shows the two steps that make complaint themes countable: consistent categorisation and deduplication. No outcome figures are published.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Complaint records with free text, product, channel, outcome and redress
- Call and chat transcripts or notes linked to the complaint
- Process maps, product terms, change and incident logs, and the control library
- A small set of complaints with expert validated root causes for testing
Systems to integrate
- Complaints and case management system
- Contact centre transcripts and conversation analytics
- Change management, incident and control library (GRC) systems
- Business intelligence and conduct risk reporting
Complexity: Medium
Clustering text is straightforward; linking themes to causes needs process, product and change data, and the value depends on a governance route that makes business owners act.
- 1
Build the evidence base first
Bring complaint text, transcripts and outcomes together for at least twelve months, with vulnerability and redress fields, before any modelling. Themes need history to show trend.
- 2
Let themes emerge, then fix the taxonomy
Run clustering on the full population, have root cause analysts name and merge themes, and publish a stable theme list. Keep complaint handling codes separate.
- 3
Link themes to owners
Map each theme to a process, product and control owner using the control library and change log, so insight goes to the person who can fix it, not back to complaints.
- 4
Validate on samples
For every theme, analysts review a sample of complaints against the AI description and record agreement. Themes below the agreement threshold are not reported.
- 5
Close the loop
Track actions to closure and measure the theme's volume and severity after each fix, and report both to the conduct or risk committee.
Guardrails
- A human root cause analyst validates every theme and cause before it is reported or actioned
- Every theme links to example complaints so a reviewer can check it
- Personal data masked in prompts; outputs report themes, not individual customers
- Vulnerability and detriment themes always escalate, regardless of volume
- Theme definitions and model changes are versioned so trends stay comparable
KPIs to instrument
- Share of complaints read by the analysis (target the full population, not a sample)
- Analyst agreement rate with AI themes and causes on samples
- Time from a theme emerging to an owner being assigned
- Complaint volume per theme before and after remediation
- Share of remediation actions with measured impact
Human in the loop
Root cause analysts validate themes and causes, business owners decide the remediation, and the conduct or risk committee reviews progress. The AI never decides redress or closes an individual complaint.
Common failure modes
- Themes that mirror the handling codes
- The model learns the existing categories and finds nothing new. Cluster on the free text and compare with codes, rather than training on codes.
- Insight with no owner
- Reports go back to the complaints team and nothing changes. Route every validated theme to a named business owner with a date.
- Plausible but wrong causes
- The AI attributes a theme to a recent change because it is in the retrieved context. Require evidence and analyst validation before a cause is reported.
- Volume blindness
- Small but severe themes, such as harm to vulnerable customers, are ranked low. Weight by severity and vulnerability, not volume alone.
What are the risks and rules?
EU AI Act
Minimal risk
Analysing complaints in aggregate to find causes is not listed in Annex III, is not a practice prohibited by Article 5 and does not decide on individuals. It does not interact with the public, so the disclosure duty in Article 50(1) does not apply; the machine readable marking of generated text in Article 50(2) is a duty of the provider of the generative model or system that writes the summaries. If the same system decided individual complaint outcomes or redress, or its themes were used to evaluate the performance of individual complaint handlers (Annex III point 4), that design would need its own assessment.
Rules that apply
Guidance
- DISP 1.3 Complaints handling rules (Financial Conduct Authority, Europe). Firms must put management controls in place to identify and remedy any recurring or systemic problems found in complaints.
- Complaints and root cause analysis: good practice and areas for improvement (Financial Conduct Authority, Europe). Thematic review of 40 firms with examples of good practice, including involving the owner of the process in root cause work and measuring the impact of fixes.
- RG 271 Internal dispute resolution (Australian Securities and Investments Commission, Asia Pacific). ASIC's standards for internal dispute resolution by Australian financial firms. Enforceable paragraph RG 271.120 requires firms to regularly analyse complaint data sets to identify systemic issues and escalate them for investigation and action.
Controls to put in place
- Documented method for theme detection and cause attribution, owned by the complaints or conduct function
- Sample based validation with recorded agreement rates per theme
- Action log from theme to owner to fix to measured impact
- Regular reporting of themes and actions to senior management and the board
- Data protection impact assessment for the use of complaint text and transcripts
Frequently asked questions
- How is this different from AI complaint handling?
- Complaint handling works one complaint at a time: recognise, investigate, respond. Root cause analysis works across all complaints to find the recurring causes behind them and get them fixed, which is what rules such as the FCA's DISP 1.3 require on top of good handling.
- Can AI find root causes on its own?
- It can read every complaint rather than a sample, cluster themes and propose likely causes. A human analyst should validate each cause, because a plausible explanation from retrieved context is not proof, and a business owner has to decide the fix.
- Who already does this?
- The public records on this page come from US agencies. The Federal Reserve Board has applied topic modelling to consumer complaint narratives from the CFPB database since 2019, and the Federal Trade Commission has classified the complaints it receives and grouped duplicates with machine learning since 2019. Of these, only the US Centers for Medicare and Medicaid Services goes as far as root causes, and it is still a pilot that finds root causes and trends in complaint cases for expert validation.
How to cite this page
Blits.ai AI Use Case Library, "AI for complaints root cause and systemic issue analysis", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/complaints-root-cause-analysis. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published