What problem does it solve?
Traditional contact centre QA reviews a tiny sample: a few calls per agent per month, scored by hand. That sample is too small to find systematic problems, too late to coach while the call is remembered, and inconsistent between reviewers. For regulated firms it is also weak evidence: when a supervisor asks whether required disclosures were given, or whether vulnerable customers were treated fairly, a small sample is not much of an answer.
Speech analytics and language models make it possible to evaluate every interaction against the same rubric within a day. The value is coverage and consistency: finding the missed disclosure, the mis sold product or the recurring complaint driver, and coaching on patterns rather than anecdotes. The risk is treating a model score as a verdict on a person, which is both unfair and, in the EU, a high risk use of AI.
How does it work?
- Capture and transcribe. Calls are transcribed after the fact (or in near real time) with speaker separation; chat and email are ingested as text. Card data is redacted.
- Score against the rubric. Each interaction is checked against configurable criteria: required disclosures, identity checks, script steps, prohibited statements, complaint and vulnerability indicators, and service behaviours.
- Flag for review. Interactions with likely breaches, complaints or vulnerability signals go to a QA or compliance reviewer with the relevant excerpt, not a bare score.
- Calibrate against humans. Reviewers regularly score the same interactions as the model; disagreements tune the rubric and prompts.
- Coach on patterns. Team leaders see recurring gaps per team and topic and coach from real examples; agents can see and dispute their own results.
- Report oversight evidence. Compliance gets coverage statistics, breach rates and trends for conduct reporting and root cause analysis.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Early adopters
- Channels
- Phone and voice, Web chat, Agent desktop
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Interactions handled | Not pooled | 167,000 | 1 | 1 vendor |
| Quality score uplift | Too few to pool | about 10% | 1 | 1 vendor |
| Handling time reduction | Too few to pool | Not pooled: up to 5% | 0plus 1 up to | 1 vendor |
Value drivers: Compliance quality, Risk and loss reduction, Customer experience, Employee productivity.
Indicative value
A contact centre with 500 agents and a QA team of 10 to 20 analysts
USD 150,000 to USD 700,000
QA analyst capacity redirected per year
How this is calculated
Formula: qaAnalysts * analystCost * timeFreed. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| QA analysts qaAnalysts, analysts | 10 | 20 | Editorial assumption of one analyst per 25 to 50 agents. Replace with your own team size. |
| Fully loaded cost per QA analyst analystCost, USD per analyst per year | 50,000 | 70,000 | Editorial assumption. Replace with your own cost. |
| Share of analyst time moved from listening and scoring to coaching and root cause work timeFreed, fraction of analyst time | 0.3 | 0.5 | Editorial assumption. Automated scoring replaces most manual listening, but calibration and review of flagged interactions remain. |
What it leaves out: Covers QA effort only. It leaves out the main value, which is finding compliance breaches, mis selling and complaint drivers that a sample misses, and the platform and transcription costs.
Who already uses it?
5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
British Gas
United Kingdom · Energy and utilities · 2026
British Gas, which runs a 20 million contact operation across voice and webchat, used to assess quality and regulatory compliance manually, which limited volume and consistency. With CallMiner it now runs millions of automated assessments across voice and webchat against its Five Steps to Customer Excellence framework and its Ofgem regulatory checks, and feeds the results into agent coaching. The automated scores had to reach at least 80% agreement with human reviewers before going live, and agents and team leaders can challenge scores. The vendor reports that quality scores improved by about 10% and that several regulatory scores now meet or exceed target.
- Quality score uplift: about 10%, quality scores against the Five Steps to Customer Excellence framework
"Quality scores against the Five Steps to Customer Excellence framework have improved by approximately 10%, with a general upward trend."
Claimed by: vendor
DoorDash
United States · Retail and ecommerce · 2026
DoorDash moved from manually reviewing a small sample of support interactions to automated evaluation of nearly all of them, reaching nearly 100% automated quality coverage across 19,000 frontline teammates in its own teams and BPO partners. It uses sentiment, comprehension and behavioural signals rather than only binary compliance checklists, and coaches from AI generated insights. Emerging problems that took days or weeks to surface are now seen in near real time.
No outcome disclosed.
Central Bank
United States · Banking · 2024
Central Bank, a group of community banks serving several states, mostly in the Midwest, runs a customer service centre with more than 3,000 interactions a day. It replaced manual QA sampling with Observe.AI's automated QA, searchable transcripts and AI tagging of call reasons, which also replaced manual disposition codes. The team went from evaluating a handful of calls a month to evaluating every call, and used the insight on agent behaviours to lower handling time.
- Interactions handled: 167,000, third quarter of 2024
"In the third quarter of 2024, the team evaluated 167,000 calls, a jump from just eight per month or 24 per quarter before adopting Auto QA."
Claimed by: vendor - Handling time reduction: up to 5%
"Identifying key agent behaviors has helped the CSC team develop data-driven plans and reduce average handle time by up to 5%, resulting in significant efficiency gains."
Claimed by: vendor
Oportun
United States · Banking · 2024
Oportun, a US consumer lender, replaced manual, sample based QA with Cresta's AI quality management across all calls, combined with real time guidance for agents. Coaching now focuses on the behaviours that drive performance, visible across every call, instead of a small sample reviewed weeks later. No quantified outcome is published.
No outcome disclosed.
VitalityHealth
United Kingdom · Insurance · 2020
VitalityHealth, a UK health insurer with more than 550 customer service advisors and over 1 million calls a year, moved from manual to automated quality assurance with CallMiner, run as a managed service by Davies Consulting. Every call is analysed and assessed in three areas: regulatory, service excellence (tone, empathy, how the call opened and closed) and process assurance, and results reach the coaching system within 24 hours. No quantified outcome is published.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- The QA rubric and regulatory disclosure requirements per product and journey
- A calibration set of interactions scored by experienced reviewers
- Access to recordings and chat logs with metadata (agent, team, product, outcome)
Systems to integrate
- Call recording and telephony platform
- Chat and messaging platforms
- QA, coaching and workforce management tools
- Complaints and conduct risk reporting
Complexity: Medium
Transcription and scoring are mature. The effort is in writing a rubric specific enough for a machine to apply, calibrating it against human reviewers, integrating recordings from the telephony platform, and agreeing with HR and employee representatives how results may be used.
- 1
Rewrite the rubric for machines
Turn vague criteria ("showed empathy") into observable checks ("acknowledged the problem before offering a solution"), and list every mandatory disclosure per journey.
- 2
Calibrate before you publish scores
Have experienced reviewers score a few hundred interactions and compare with the model. Publish only criteria where agreement is at least as good as between two humans. British Gas tested until its automated scores reached at least 80% agreement with human reviewers before going live.
- 3
Start with compliance checks
Mandatory disclosures and prohibited statements are the clearest criteria and the strongest oversight evidence. Add softer service behaviours later.
- 4
Agree the rules of use
Decide with HR, legal and employee representatives how results feed coaching and whether they may affect evaluation, pay or discipline, and document the assessment.
- 5
Give agents visibility and a dispute route
Let agents see their scored interactions and challenge them. Disputes are also a calibration signal.
- 6
Close the loop to root causes
Feed recurring breaches and complaint drivers to product, process and training owners, not only to individual coaching.
Guardrails
- Scores are inputs for human review, never automatic sanctions
- Rubric criteria published only after calibration against human reviewers
- No inference of agents' emotions
- Card data and special category data redacted before scoring and storage
- Agents can see and dispute their results
KPIs to instrument
- Share of interactions evaluated automatically
- Agreement between model and human reviewers per criterion
- Confirmed breach rate for mandatory disclosures, per journey
- Disputed scores and their outcome
- Complaint and repeat contact trends after coaching
Human in the loop
QA and compliance reviewers confirm every flagged breach before it is recorded or acted on. Team leaders decide on coaching, and any consequence for an individual follows the normal HR process with human judgment. Reviewers recalibrate the rubric at least quarterly.
Common failure modes
- Scores treated as facts
- An uncalibrated score drives performance ratings. Calibrate per criterion and keep humans in every consequential decision.
- Rubric too vague for a model
- Criteria such as "professional tone" give noisy scores. Rewrite into observable behaviours.
- Surveillance backlash
- Agents experience total monitoring without transparency, and trust and retention fall. Be open about what is measured and let agents dispute.
- Finding problems nobody fixes
- Breaches are counted but root causes in products or processes remain. Route themes to owners with deadlines.
What are the risks and rules?
EU AI Act
High risk
Scoring individual agents' interactions to monitor and evaluate their performance and behaviour falls under Annex III point 4(b), employment and worker management. The Article 6(3) exception does not apply where the system profiles natural persons. Inferring agents' emotions is prohibited under Article 5(1)(f), except for medical or safety reasons. Inferring customers' emotions from their voice is emotion recognition on biometric data: high risk under Annex III point 1(c), and Article 50(3) requires informing the people exposed to it. Analytics that only aggregate interaction themes without evaluating individuals can fall outside the high risk category.
Rules that apply
Guidance
- Annex III, high risk AI systems referred to in Article 6(2) (European Union, Europe). Point 4(b) covers AI systems intended to monitor and evaluate the performance and behaviour of workers.
- Article 5, prohibited AI practices (European Union, Europe). Point 1(f) prohibits inferring emotions of natural persons in the workplace, except for medical or safety reasons.
- Article 50, transparency obligations for providers and deployers of certain AI systems (European Union, Europe). Point 3 requires deployers of an emotion recognition system to inform the people exposed to it, which matters when voice analytics infers customer sentiment from speech.
- Employment practices and data protection: monitoring workers (UK Information Commissioner's Office, Europe). Guidance on transparency, necessity and impact assessments when monitoring workers, including through automated tools. The ICO marks it as under review after the Data (Use and Access) Act.
Controls to put in place
- Data protection impact assessment and, in the EU, the high risk obligations for the deployer
- Documented calibration results per criterion and per model version
- Written rules on how scores may and may not be used for individual decisions
- Agent access to their results and a dispute process
- Retention limits and access controls on recordings and transcripts
Frequently asked questions
- Can AI really review every call?
- Yes, coverage is the main change. Observe.AI reports that Central Bank, a US community bank group, evaluated 167,000 calls in the third quarter of 2024, up from 24 per quarter before automated QA, and that DoorDash reached nearly 100% automated quality coverage across 19,000 agents.
- Is automated agent scoring high risk under the EU AI Act?
- When it evaluates individual workers, yes: Annex III point 4(b) covers monitoring and evaluating the performance and behaviour of workers. Inferring agents' emotions from biometric data such as voice is prohibited under Article 5(1)(f), except for medical or safety reasons. Aggregate analytics on themes, without scoring individuals, carry less risk.
- How do we know the scores are right?
- Calibrate. Have experienced reviewers score the same interactions as the model, publish only criteria where agreement is good enough, and repeat after every rubric or model change. British Gas, for example, went live only after its automated scores reached at least 80% agreement with human reviewers, and lets agents and team leaders challenge scores.
How to cite this page
Blits.ai AI Use Case Library, "AI quality and compliance monitoring of every customer interaction", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/call-quality-and-compliance-monitoring. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published