AI use case

AI quality and compliance monitoring of every customer interaction

Automated quality assurance that transcribes and scores every customer interaction, voice and chat, against the organization's own rubric, checking required disclosures and script adherence, flagging conduct and mis selling risk, and surfacing coaching opportunities, instead of the small sample a human QA team can review.

By Len Debets · Last verified 27 September 2026 · 5 public deployments

About 10%
Reported quality score uplift
British Gas, vendor claim.
Up to 5%
Reported handling time reduction
Central Bank, vendor claim.
USD 150,000 to USD 700,000
Indicative value per year
A contact centre with 500 agents and a QA team of 10 to 20 analysts. Worked example, see how it is calculated.

What problem does it solve?

Traditional contact centre QA reviews a tiny sample: a few calls per agent per month, scored by hand. That sample is too small to find systematic problems, too late to coach while the call is remembered, and inconsistent between reviewers. For regulated firms it is also weak evidence: when a supervisor asks whether required disclosures were given, or whether vulnerable customers were treated fairly, a small sample is not much of an answer.

Speech analytics and language models make it possible to evaluate every interaction against the same rubric within a day. The value is coverage and consistency: finding the missed disclosure, the mis sold product or the recurring complaint driver, and coaching on patterns rather than anecdotes. The risk is treating a model score as a verdict on a person, which is both unfair and, in the EU, a high risk use of AI.

How does it work?

  1. Capture and transcribe. Calls are transcribed after the fact (or in near real time) with speaker separation; chat and email are ingested as text. Card data is redacted.
  2. Score against the rubric. Each interaction is checked against configurable criteria: required disclosures, identity checks, script steps, prohibited statements, complaint and vulnerability indicators, and service behaviours.
  3. Flag for review. Interactions with likely breaches, complaints or vulnerability signals go to a QA or compliance reviewer with the relevant excerpt, not a bare score.
  4. Calibrate against humans. Reviewers regularly score the same interactions as the model; disagreements tune the rubric and prompts.
  5. Coach on patterns. Team leaders see recurring gaps per team and topic and coach from real examples; agents can see and dispute their own results.
  6. Report oversight evidence. Compliance gets coverage statistics, breach rates and trends for conduct reporting and root cause analysis.
Audience
Back office
Autonomy
Supervised agent
Adoption
Early adopters
Channels
Phone and voice, Web chat, Agent desktop

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI quality and compliance monitoring of every customer interaction
KPIMedianReported rangeData pointsClaimed by
Interactions handledNot pooled
167,000
11 vendor
Quality score upliftToo few to pool
about 10%
11 vendor
Handling time reductionToo few to pool
Not pooled: up to 5%
0plus 1 up to1 vendor

Value drivers: Compliance quality, Risk and loss reduction, Customer experience, Employee productivity.

Indicative value

A contact centre with 500 agents and a QA team of 10 to 20 analysts

USD 150,000 to USD 700,000

QA analyst capacity redirected per year

How this is calculated

Formula: qaAnalysts * analystCost * timeFreed. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
QA analysts qaAnalysts, analysts1020Editorial assumption of one analyst per 25 to 50 agents. Replace with your own team size.
Fully loaded cost per QA analyst analystCost, USD per analyst per year50,00070,000Editorial assumption. Replace with your own cost.
Share of analyst time moved from listening and scoring to coaching and root cause work timeFreed, fraction of analyst time0.30.5Editorial assumption. Automated scoring replaces most manual listening, but calibration and review of flagged interactions remain.

What it leaves out: Covers QA effort only. It leaves out the main value, which is finding compliance breaches, mis selling and complaint drivers that a sample misses, and the platform and transcription costs.

Who already uses it?

5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

British Gas

United Kingdom · Energy and utilities · 2026

ScaledGrade C

British Gas, which runs a 20 million contact operation across voice and webchat, used to assess quality and regulatory compliance manually, which limited volume and consistency. With CallMiner it now runs millions of automated assessments across voice and webchat against its Five Steps to Customer Excellence framework and its Ofgem regulatory checks, and feeds the results into agent coaching. The automated scores had to reach at least 80% agreement with human reviewers before going live, and agents and team leaders can challenge scores. The vendor reports that quality scores improved by about 10% and that several regulatory scores now meet or exceed target.

  • Quality score uplift: about 10%, quality scores against the Five Steps to Customer Excellence framework
    "Quality scores against the Five Steps to Customer Excellence framework have improved by approximately 10%, with a general upward trend."
    Claimed by: vendor

DoorDash

United States · Retail and ecommerce · 2026

ScaledGrade C

DoorDash moved from manually reviewing a small sample of support interactions to automated evaluation of nearly all of them, reaching nearly 100% automated quality coverage across 19,000 frontline teammates in its own teams and BPO partners. It uses sentiment, comprehension and behavioural signals rather than only binary compliance checklists, and coaches from AI generated insights. Emerging problems that took days or weeks to surface are now seen in near real time.

No outcome disclosed.

Central Bank

United States · Banking · 2024

ScaledGrade C

Central Bank, a group of community banks serving several states, mostly in the Midwest, runs a customer service centre with more than 3,000 interactions a day. It replaced manual QA sampling with Observe.AI's automated QA, searchable transcripts and AI tagging of call reasons, which also replaced manual disposition codes. The team went from evaluating a handful of calls a month to evaluating every call, and used the insight on agent behaviours to lower handling time.

  • Interactions handled: 167,000, third quarter of 2024
    "In the third quarter of 2024, the team evaluated 167,000 calls, a jump from just eight per month or 24 per quarter before adopting Auto QA."
    Claimed by: vendor
  • Handling time reduction: up to 5%
    "Identifying key agent behaviors has helped the CSC team develop data-driven plans and reduce average handle time by up to 5%, resulting in significant efficiency gains."
    Claimed by: vendor

Oportun

United States · Banking · 2024

ScaledGrade C

Oportun, a US consumer lender, replaced manual, sample based QA with Cresta's AI quality management across all calls, combined with real time guidance for agents. Coaching now focuses on the behaviours that drive performance, visible across every call, instead of a small sample reviewed weeks later. No quantified outcome is published.

No outcome disclosed.

VitalityHealth

United Kingdom · Insurance · 2020

ScaledGrade C

VitalityHealth, a UK health insurer with more than 550 customer service advisors and over 1 million calls a year, moved from manual to automated quality assurance with CallMiner, run as a managed service by Davies Consulting. Every call is analysed and assessed in three areas: regulatory, service excellence (tone, empathy, how the call opened and closed) and process assurance, and results reach the coaching system within 24 hours. No quantified outcome is published.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • The QA rubric and regulatory disclosure requirements per product and journey
  • A calibration set of interactions scored by experienced reviewers
  • Access to recordings and chat logs with metadata (agent, team, product, outcome)

Systems to integrate

  • Call recording and telephony platform
  • Chat and messaging platforms
  • QA, coaching and workforce management tools
  • Complaints and conduct risk reporting

Complexity: Medium

Transcription and scoring are mature. The effort is in writing a rubric specific enough for a machine to apply, calibrating it against human reviewers, integrating recordings from the telephony platform, and agreeing with HR and employee representatives how results may be used.

  1. 1

    Rewrite the rubric for machines

    Turn vague criteria ("showed empathy") into observable checks ("acknowledged the problem before offering a solution"), and list every mandatory disclosure per journey.

  2. 2

    Calibrate before you publish scores

    Have experienced reviewers score a few hundred interactions and compare with the model. Publish only criteria where agreement is at least as good as between two humans. British Gas tested until its automated scores reached at least 80% agreement with human reviewers before going live.

  3. 3

    Start with compliance checks

    Mandatory disclosures and prohibited statements are the clearest criteria and the strongest oversight evidence. Add softer service behaviours later.

  4. 4

    Agree the rules of use

    Decide with HR, legal and employee representatives how results feed coaching and whether they may affect evaluation, pay or discipline, and document the assessment.

  5. 5

    Give agents visibility and a dispute route

    Let agents see their scored interactions and challenge them. Disputes are also a calibration signal.

  6. 6

    Close the loop to root causes

    Feed recurring breaches and complaint drivers to product, process and training owners, not only to individual coaching.

Guardrails

  • Scores are inputs for human review, never automatic sanctions
  • Rubric criteria published only after calibration against human reviewers
  • No inference of agents' emotions
  • Card data and special category data redacted before scoring and storage
  • Agents can see and dispute their results

KPIs to instrument

  • Share of interactions evaluated automatically
  • Agreement between model and human reviewers per criterion
  • Confirmed breach rate for mandatory disclosures, per journey
  • Disputed scores and their outcome
  • Complaint and repeat contact trends after coaching

Human in the loop

QA and compliance reviewers confirm every flagged breach before it is recorded or acted on. Team leaders decide on coaching, and any consequence for an individual follows the normal HR process with human judgment. Reviewers recalibrate the rubric at least quarterly.

Common failure modes

Scores treated as facts
An uncalibrated score drives performance ratings. Calibrate per criterion and keep humans in every consequential decision.
Rubric too vague for a model
Criteria such as "professional tone" give noisy scores. Rewrite into observable behaviours.
Surveillance backlash
Agents experience total monitoring without transparency, and trust and retention fall. Be open about what is measured and let agents dispute.
Finding problems nobody fixes
Breaches are counted but root causes in products or processes remain. Route themes to owners with deadlines.

What are the risks and rules?

EU AI Act

High risk

Scoring individual agents' interactions to monitor and evaluate their performance and behaviour falls under Annex III point 4(b), employment and worker management. The Article 6(3) exception does not apply where the system profiles natural persons. Inferring agents' emotions is prohibited under Article 5(1)(f), except for medical or safety reasons. Inferring customers' emotions from their voice is emotion recognition on biometric data: high risk under Annex III point 1(c), and Article 50(3) requires informing the people exposed to it. Analytics that only aggregate interaction themes without evaluating individuals can fall outside the high risk category.

Guidance

Controls to put in place

  • Data protection impact assessment and, in the EU, the high risk obligations for the deployer
  • Documented calibration results per criterion and per model version
  • Written rules on how scores may and may not be used for individual decisions
  • Agent access to their results and a dispute process
  • Retention limits and access controls on recordings and transcripts

Frequently asked questions

Can AI really review every call?
Yes, coverage is the main change. Observe.AI reports that Central Bank, a US community bank group, evaluated 167,000 calls in the third quarter of 2024, up from 24 per quarter before automated QA, and that DoorDash reached nearly 100% automated quality coverage across 19,000 agents.
Is automated agent scoring high risk under the EU AI Act?
When it evaluates individual workers, yes: Annex III point 4(b) covers monitoring and evaluating the performance and behaviour of workers. Inferring agents' emotions from biometric data such as voice is prohibited under Article 5(1)(f), except for medical or safety reasons. Aggregate analytics on themes, without scoring individuals, carry less risk.
How do we know the scores are right?
Calibrate. Have experienced reviewers score the same interactions as the model, publish only criteria where agreement is good enough, and repeat after every rubric or model change. British Gas, for example, went live only after its automated scores reached at least 80% agreement with human reviewers, and lets agents and team leaders challenge scores.

How to cite this page

Blits.ai AI Use Case Library, "AI quality and compliance monitoring of every customer interaction", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/call-quality-and-compliance-monitoring. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryBanking

Real time AI assist for contact centre agents

A real time copilot for human contact centre agents during a live call or chat: it transcribes the conversation as it happens, surfaces the relevant knowledge and next step, drafts responses, and writes the after call summary and CRM notes, while the agent stays in control of what is said and done.

Deployments
5 public, best grade B
Reported productivity gain
15%
Definity, vendor claim
Cross industryBanking

AI roleplay training for customer conversations

A training simulator in which generative AI plays a realistic customer, by voice or text, so service, sales and crisis staff can rehearse difficult conversations as often as they need before they handle live ones, and receive structured feedback against the organization's own standards.

Deployments
3 public, best grade B
Reported conversion uplift
21%
GoHealth, vendor claim
Cross industryBanking

AI for complaints root cause and systemic issue analysis

AI that reads the free text of complaints across all channels, clusters them into themes, separates systemic causes from one off events, links each theme to the product, process or control behind it and routes the insight to the owner who can fix it, with a human validating every root cause and every remediation.

Deployments
3 public, best grade B
Autonomy
Copilot
Cross industryTelecommunications

AI sales call coaching and CRM update

AI for sales teams that analyses sales calls and meetings against the team's own sales method to coach sellers and their managers, and writes the call summary, next steps and opportunity updates into the CRM for the seller to confirm. Its purpose is winning deals and building selling skill, not the regulated advice record or general meeting notes.

Deployments
4 public, best grade C
Reported time saved per task
3 minutes
Sandvik Coromant, organization claim
Cross industryRetail and ecommerce

AI for voice of the customer and feedback analysis

AI that reads every piece of free text customer feedback, such as survey verbatims, NPS comments, reviews, social posts, chat and call transcripts, and turns it into themes, sentiment, drivers and suggested actions that a named owner can act on, so the organization hears all of its customers instead of a sample.

Deployments
5 public, best grade B
Reported accuracy
84%
SBF Group, vendor claim