What problem does it solve?
In regulated industries a complaint is not whatever the customer calls a complaint. Financial regulators define it broadly (any expression of dissatisfaction, about the firm's service or products), and a complaint made in a phone call or a chat counts as much as a letter. Firms must acknowledge quickly, resolve within set deadlines, explain the outcome in writing and report volumes and root causes. Missing one is a breach of the rules, not just a service failure.
Much of the handler's work sits around the decision: reading the history across systems, pulling statements and call notes, working out what went wrong, and writing a response that is accurate, fair and clear. Meanwhile complaints hidden in ordinary conversations are never logged, so they are never fixed. Customers now also use generative AI to write complaints, which the UK Financial Ombudsman Service says can produce long, unfocused submissions with fabricated laws, misquoted regulations or invented past decisions.
How does it work?
- Recognize. Every channel (calls, chats, emails, letters, social) is screened for expressions of dissatisfaction against the regulatory definition, so a complaint inside an ordinary call is flagged instead of lost.
- Log and classify. The agent opens the case, sets the product, root cause and severity, flags vulnerability and possible systemic issues, and starts the deadline clock.
- Investigate. It gathers the evidence from the relevant systems (transactions, call notes, previous contacts, policies in force at the time) and writes a summary of what happened.
- Draft. It drafts the acknowledgement and the final response from approved wording, with the reasoning and the redress calculation shown separately for the handler.
- Decide and send, by a human. A complaint handler reviews the evidence, decides the outcome and approves the letter. The agent tracks deadlines and feeds root causes to the teams that can fix them.
- Audience
- Employee facing
- Autonomy
- Copilot
- Adoption
- Early adopters
- Channels
- Agent desktop, Internal tools, Email, Web chat, Phone and voice
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Time saved per task | Too few to pool | about 5 minutes | 1 | 1 organization |
Value drivers: Compliance quality, Employee productivity, Speed and cycle time, Customer experience.
Indicative value
A retail bank that handles 50,000 complaints a year
USD 166,667 to USD 1.2 million
Complaint handler time released per year
How this is calculated
Formula: complaints * minutesSaved / 60 * costPerHour. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Complaints handled per year complaints, complaints per year | 50,000 | 50,000 | The reference bank. |
| Handler minutes saved per complaint on classification, evidence gathering and drafting minutesSaved, minutes per complaint | 5 | 20 | The low end is the only reported figure on this page, which covers classification alone (Lloyds Banking Group reports complaint classification in 1 second instead of about 5 minutes). The high end adds evidence gathering and drafting, for which no deployment has disclosed a figure; editorial assumption, replace with your own. |
| Fully loaded cost of a complaint handler hour costPerHour, USD per hour | 40 | 70 | Editorial assumption, replace with your own cost. |
What it leaves out: Handler time only. It leaves out lower redress and ombudsman fees from better first responses, fewer missed deadlines, the value of fixing root causes earlier and the cost of the AI and integrations.
Who already uses it?
2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
NatWest Group
United Kingdom · Banking · 2026
NatWest is testing an agentic AI system that investigates customer complaints across several data sources and presents a summarised view to a complaint handler, who approves it, with the aim of speeding up handling and resolution. The trial runs in the FCA's AI Live Testing environment, with performance tracked daily on task accuracy, coherence and hallucination, and every AI generated summary subject to human oversight. The bank plans to use production like environments before any move to live use. It is a pilot; no outcome figures are disclosed.
No outcome disclosed.
Lloyds Banking Group
United Kingdom · Banking · 2025
Lloyds Banking Group lists complaints handling and automation among the roughly 50 generative AI use cases it had live in 2025. In its 2025 results presentation the bank reports that complaint classification now takes 1 second instead of about 5 minutes. The bank attributes about GBP 50 million of P&L benefit in 2025 to its generative AI use cases as a whole and does not break out the share of the complaints use case.
- Time saved per task: about 5 minutes, in 2025
"Outcome: Classification times reduced to 1 second (from c.5 mins)"
Claimed by: organization
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- The regulatory complaint definition and internal taxonomy of products, causes and severities
- Historical complaints with outcomes and final response letters
- Approved response templates, clauses and redress rules
- Access to call transcripts and chat logs for recognition
Systems to integrate
- Complaint or case management system
- Contact centre transcripts and chat platforms
- Core banking, card and product systems for evidence
- Document generation and correspondence
- Management information and regulatory reporting
Complexity: Medium
Complaint classification is live at Lloyds Banking Group and investigation summaries are in testing at NatWest. The difficulty is reaching the evidence across many systems, keeping drafts factually right, recognizing complaints in unstructured conversations and proving to the regulator that nothing is missed.
- 1
Start with recognition and logging
Screen transcripts and messages for complaints the contact centre did not log, and measure how many are found. This reduces conduct risk before any drafting is automated.
- 2
Summarise the evidence for handlers
Build the investigation summary next, with links to every source, and measure how often handlers agree with it. NatWest's pilot follows this pattern: the agent investigates across data sources and presents a summary to a handler for approval.
- 3
Draft responses from approved wording
Generate drafts from templates and clauses, with the decision and any redress set by the handler, and track edit rates per section.
- 4
Close the loop on root causes
Aggregate causes across cases weekly and route them to product and process owners; spot systemic issues that affect many customers early.
- 5
Test with your regulator in mind
Keep an evaluation set of real cases scored for task accuracy and hallucination, tracked over time. NatWest runs its trial inside the FCA's AI Live Testing with daily tracking of such metrics.
Guardrails
- Every outcome and every response letter is approved by a human complaint handler
- The agent may escalate a case to a complaint but never downgrade a flagged complaint on its own
- Drafts cite the evidence they rely on; unsupported statements are flagged for the handler
- Deadlines are computed by rules and alerted, never estimated by the model
- Vulnerability and systemic issue flags route to specialist teams
KPIs to instrument
- Complaints recognized in conversations that were not logged manually
- Time from receipt to acknowledgement and to final response
- Handler agreement with classifications and edit rate on drafts
- Deadline breaches and ombudsman referral and overturn rates
- Root causes identified and fixed
Human in the loop
Complaint handlers decide every outcome and approve every letter. Quality assurance samples AI classifications and drafts each week, and a senior owner signs off the recognition rules and thresholds, because a missed complaint is a conduct failure.
Common failure modes
- Containment over recognition
- A customer facing assistant tuned for containment answers a complaint as a question and never logs it. Screen every conversation against the complaint definition.
- Plausible but wrong responses
- A draft misstates facts or policy and the handler, under time pressure, sends it. Show evidence next to every claim and track edit rates.
- Automated unfairness
- Triage deprioritises complex or vulnerable cases. Test routing outcomes by customer group and keep humans on vulnerability.
- AI written complaints overwhelm triage
- Long AI drafted submissions with invented legal references slow handling. Summarise the customer's actual points and check references before responding.
What are the risks and rules?
EU AI Act
Depends on design
Complaint handling is not listed in Annex III, so internal classification and drafting for a handler who decides is minimal risk. Where the agent talks to customers to take the complaint, Article 50(1) requires telling them they are dealing with AI. Only a system that also assessed creditworthiness or priced life and health insurance (Annex III point 5(b) or 5(c)) would be high risk for that part.
Rules that apply
Guidance
- DISP 1: Treating complainants fairly (Financial Conduct Authority, Europe). The UK rules for complaint handling by financial firms, including prompt acknowledgement, the eight week time limit, final response requirements and complaint reporting.
- RG 271 Internal dispute resolution (Australian Securities and Investments Commission, Asia Pacific). ASIC's standards for internal dispute resolution by Australian financial firms, including what counts as a complaint and maximum response timeframes.
- Embracing AI's transformational impact on consumer complaints (Financial Ombudsman Service, Europe). The UK ombudsman's view of how AI is changing complaints, including consumers' use of generative AI and firms' automated triage.
Controls to put in place
- Recognition rules mapped to the regulatory complaint definition, owned by a senior manager
- Human approval of every outcome and final response, recorded in the case
- Audit trail of classifications, evidence used and draft changes
- Regular accuracy and hallucination testing on real cases
- Root cause and systemic issue reporting to governance forums
Frequently asked questions
- Can AI decide complaint outcomes?
- It should not. NatWest's agent investigates and presents a summarised view to a complaint handler for approval, and the bank says all AI generated summaries are subject to strict human oversight. The other reported use on this page, at Lloyds Banking Group, is classifying complaints.
- Where does AI save the most time in complaints?
- Among the deployments on this page, the only reported gain is in classification: Lloyds Banking Group reports that complaint classification takes 1 second instead of about 5 minutes. NatWest is testing an agent that gathers evidence from several data sources into one summary for the handler, but has not disclosed results.
- What is the biggest risk?
- A complaint that is never recognized, for example because a customer facing bot treats it as a question to contain. Rules such as the FCA's DISP 1 require complaints to be recorded and resolved within set time limits, a prompt written acknowledgement and a final response within eight weeks, unless the complaint is resolved by the third business day, so recognition deserves as much testing as drafting.
How to cite this page
Blits.ai AI Use Case Library, "AI agent for complaints recognition, investigation and response", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/complaints-handling-agent. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published