What problem does it solve?
New contact centre, sales and crisis line staff learn the hardest conversations on real customers: the angry caller, the fraud victim, the customer in financial hardship, the person in crisis. Classroom roleplay with colleagues or trainers is limited by trainer time, feels artificial and rarely covers the full range of situations, so new hires often reach the floor with little practice. The risk is long ramp up times, inconsistent handling of disclosures and vulnerability, and avoidable harm to the first customers each new hire serves.
Regulated firms have an extra reason to care. Conduct rules expect staff to recognise vulnerability, give required disclosures and treat customers fairly; in the UK, the FCA's guidance on vulnerable customers asks firms to ensure frontline staff have the skills and capability to recognise and respond to vulnerability. Trainer led roleplay leaves little evidence of what was practised and how well.
How does it work?
- Build scenarios from real work. Training and quality teams write scenarios from real, anonymized contact reasons: a disputed charge, a lost card abroad, a hardship request, a complaint, a sales conversation with required disclosures. Each scenario has a persona, a goal, facts the trainee must find out and behaviours to test.
- The AI plays the customer. A model plays the persona by voice or text, reacts to what the trainee says, becomes calmer or more upset depending on how the conversation goes, and raises the objections or cues the scenario calls for.
- Score against the rubric. After the conversation the AI scores the transcript against the organization's rubric (verification steps, required disclosures, empathy, accuracy of information, next steps) and quotes the moments behind each score.
- Give targeted feedback and repeat. The trainee gets specific feedback and can retry the same scenario or a harder variant immediately.
- Report to trainers, not to discipline. Trainers see progress per skill and per cohort and spend their time coaching where the simulator shows gaps.
- Audience
- Employee facing
- Autonomy
- Assist
- Adoption
- Early adopters
- Channels
- Internal tools, Phone and voice, Web chat
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Conversion uplift | Too few to pool | 21% | 1 | 1 vendor |
| Interactions handled | Not pooled | at least 1 million | 1 | 1 organization |
| Time to proficiency reduction | Too few to pool | 55% | 1 | 1 vendor |
Value drivers: Employee productivity, Customer experience, Compliance quality, Speed and cycle time.
Indicative value
A contact centre that hires 200 new agents a year
USD 32,400 to USD 450,000
Value of productive time gained by faster ramp up per year
How this is calculated
Formula: hires * rampWeeks * rampReduction * productivityGap * weeklyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| New agents trained per year hires, agents per year | 200 | 200 | The reference organization. Replace with your own hiring volume. |
| Weeks from start to full proficiency today rampWeeks, weeks | 6 | 10 | Editorial assumption. Replace with your own ramp time. |
| Share of ramp time removed by simulated practice rampReduction, fraction of ramp time | 0.1 | 0.3 | Conservative against the evidence on this page (GoHealth's vendor reports onboarding cut from nine weeks to four, a 55% saving), because that figure is a single vendor reported case. |
| Productivity shortfall of a new agent during ramp up productivityGap, fraction of a fully proficient agent | 0.3 | 0.5 | Editorial assumption. |
| Fully loaded weekly cost of an agent weeklyCost, USD per week | 900 | 1,500 | Editorial assumption. Replace with your own cost. |
What it leaves out: Ramp up value only. It leaves out trainer time saved, lower early attrition, fewer complaints and conduct breaches from new hires, and the licence and scenario authoring costs of the simulator.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Bank of America
United States · Banking · 2024
The Academy, Bank of America's onboarding, education and professional development organization, uses AI conversation simulators in which employees practise different types of client interactions and receive real time feedback. The bank reports more than one million simulations completed in 2024 and says many employees note that practising client conversations helps them deliver better and more consistent service. No proficiency or client outcome figures are published.
- Interactions handled: at least 1 million, simulations completed by employees in 2024
"Employees completed over 1 million simulations last year, with many noting that practicing client conversations helps them deliver better and more consistent service."
Claimed by: organization
U.S. Department of Veterans Affairs
United States · Government and public sector · 2024
The Veterans Crisis Line trains new crisis responders with ReflexAI simulations in which generative AI plays eight Veteran personas, each with its own configured motivation and crisis, so trainees can practise in a low risk setting before live calls. Each simulated call produces a scoring summary of strengths and areas of growth. The tools launched with the first full cohort of new trainees in May 2024, and VA's 2025 AI use case inventory lists the system as deployed and not high impact.
No outcome disclosed.
GoHealth
United States · Insurance · 2022
GoHealth, a health insurance marketplace focused on Medicare, uses AI role play partners so licensed benefits consultants practise sales and compliance heavy conversations, both in onboarding and in ongoing training. For new hires, practice is built into the training weeks instead of a separate three week block of practice calls. After a pilot in the fourth quarter of 2022 GoHealth signed a long term contract. The vendor reports shorter onboarding, a higher sales conversion rate in the pilot and a higher trainee to trainer ratio.
- Time to proficiency reduction: 55%
"That’s a 55% time saving that allows new hires to start making effective calls five weeks earlier than before."
Claimed by: vendor - Conversion uplift: 21%, Q4 2022 pilot, sales rates 10 days before versus 10 days after about 34 minutes of AI practice
"Average 21% increase in sales conversions after 34 minutes of practice"
Claimed by: vendor
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- The contact reason report, to choose scenarios by volume and risk
- The quality assurance rubric and required disclosures per conversation type
- Anonymized example transcripts or call recordings for realistic personas
- Vulnerability and complaint handling policies the scenarios must test
Systems to integrate
- Learning management system for assignments and completion records
- Single sign on for trainees and trainers
- Voice or telephony softphone for realistic practice calls
Complexity: Low
No integration with customer systems is needed. The effort is in writing good scenarios and rubrics with the quality and compliance teams, calibrating the scoring against human assessors and making voice latency low enough to feel like a real call.
- 1
Pick scenarios by risk and volume
Start with five to ten scenarios that new hires find hardest and that carry conduct risk, such as a hardship request, a scam victim or a complaint. Add routine ones later.
- 2
Write rubrics with quality and compliance
Use the same rubric the quality team uses on live calls, so practice and assessment measure the same things. Mark which items are mandatory, such as identity checks and disclosures.
- 3
Calibrate scoring against humans
Have experienced assessors score a sample of simulated conversations and compare with the AI scores. Fix rubric items where they disagree before trainees see scores.
- 4
Keep personas realistic, not cruel
Let the persona escalate and de escalate in response to the trainee, but keep abuse within what staff actually meet, and give trainees a way to stop a session.
- 5
Blend into the programme
Integrate practice into training weeks rather than bolting it on, as GoHealth did by folding practice calls into training, and let trainers use the reports to target coaching.
- 6
Measure on the floor
Compare ramp time, quality scores and complaint rates of trained cohorts with earlier cohorts on the same contact mix.
Guardrails
- Scores are used for practice and coaching, not for promotion, pay or termination decisions without human review
- No inference of trainees' emotions from voice or face, which the EU AI Act prohibits in the workplace outside medical or safety reasons
- Scenarios and model answers use only approved policies, products and disclosure wording
- No real customer data in personas; examples are anonymized before they become scenarios
- Trainees can see their transcripts and scores and contest a score with a trainer
KPIs to instrument
- Time to proficiency per cohort, before and after
- Practice sessions per trainee and scenario coverage
- Agreement between AI scores and human assessor scores on a sample
- Live quality scores and complaint rates in the first three months on the floor
- Trainee rating of realism and usefulness
Human in the loop
Trainers and quality leads own the scenarios and rubrics, review the AI's scoring on a sample every cohort, and make every certification or sign off decision. The AI gives practice and feedback; a human decides whether someone is ready for live customers.
Common failure modes
- Scoring that trainees do not trust
- Scores that disagree with what trainers say undermine the tool. Calibrate against human assessors and show the transcript evidence behind each score.
- Unrealistic customers
- Personas that are too easy or cartoonishly hostile teach the wrong lessons. Build personas from real contact reasons and review them with experienced agents.
- Training on outdated policy
- The simulated conversation rewards a disclosure or process that has changed. Tie scenarios to the policy owner and review them when policy changes.
- Practice data used as surveillance
- Using practice scores in performance management kills honest practice and can make the system high risk. Keep practice and performance evaluation separate by design.
What are the risks and rules?
EU AI Act
Depends on design
Used only for practice and feedback, the simulator is limited risk. Article 50 requires that people know they are interacting with AI unless that is obvious from the context, as it usually is in a training session, and the provider must mark synthetic voice or text output as AI generated in a machine readable format. It becomes high risk under Annex III point 4(b) if its scores are used to evaluate the performance of workers or to decide on their promotion or termination, and can fall under point 3(b) when a vocational training institution uses it to evaluate learning outcomes. Inferring trainees' emotions from voice or face in the workplace is prohibited under Article 5(1)(f), except for medical or safety reasons.
Rules that apply
Guidance
- Annex III: High risk AI systems referred to in Article 6(2) (European Union, Europe). Point 4(b) covers AI used to monitor and evaluate the performance and behaviour of persons in work related relationships, which is where training scores can end up. Point 3(b) covers AI that evaluates learning outcomes in educational and vocational training institutions.
- Article 5: Prohibited AI practices (European Union, Europe). Point (f) prohibits AI that infers the emotions of a natural person in the workplace or in education institutions, except for medical or safety reasons, which rules out scoring trainees' emotions from their voice or face.
- Article 50, transparency obligations for providers and deployers of certain AI systems (European Union, Europe). Requires that people are told they are interacting with AI unless it is obvious from the context, and that providers mark synthetic audio and text output as artificially generated.
- FG21/1: guidance for firms on the fair treatment of vulnerable customers (Financial Conduct Authority, Europe). Asks UK financial services firms to ensure frontline staff have the skills and capability to recognise and respond to customers in vulnerable circumstances, which scenario practice can support.
Controls to put in place
- Written purpose limitation that keeps practice scores out of performance management
- Periodic calibration of AI scores against human assessors, with results recorded
- Scenario and rubric change control owned by training and compliance
- Trainee notice of how transcripts and scores are stored, who sees them and for how long
- Inventory entry for the simulator with an accountable owner
Frequently asked questions
- Does AI roleplay actually shorten ramp up time?
- Public evidence is early and mostly vendor reported. GoHealth's vendor reports onboarding cut from nine weeks to four after folding AI practice into training, and Bank of America reports more than one million simulations completed by employees in 2024. Measure ramp time and live quality on your own cohorts before and after.
- Is AI roleplay training high risk under the EU AI Act?
- Not when it is used only for practice and feedback; then the Article 50 transparency rules apply. It becomes high risk if its scores are used to evaluate employees' performance or decide on promotion or termination (Annex III point 4(b)), and inferring trainees' emotions in the workplace is prohibited (Article 5(1)(f)).
- Can it be used for sensitive conversations such as crisis calls?
- Yes, with care. The US Department of Veterans Affairs trains new Veterans Crisis Line responders on AI simulations with eight Veteran personas, and each simulated call produces a scoring summary of strengths and areas of growth. Scenarios for sensitive topics need expert review and a way for trainees to stop a session.
How to cite this page
Blits.ai AI Use Case Library, "AI roleplay training for customer conversations", last verified 26 September 2026, https://www.blits.ai/ai-use-cases/conversation-roleplay-training. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published