AI use case

AI tutor that coaches students through problems

An AI tutor that works with a student on course material in a conversation, asking questions and giving hints instead of handing over answers, grounded in the course content and set up by the school or teacher, with limits on use and a clear route to a human teacher.

By Len Debets · Last verified 27 September 2026 · 3 public deployments

USD 15,000 to USD 840,000
Indicative value per year
A school district with 10,000 students in grades 6 to 12. Worked example, see how it is calculated.

What problem does it solve?

A 2026 working paper on AI tutoring sums up the research on human tutoring: the best one to one programmes produce gains of a third of a standard deviation or more, but high dosage tutoring often costs several thousand dollars per student per year. In large courses the queue for help is the bottleneck: Harvard's CS50 recalls times when office hours became unmanageable and the average wait could be as long as an hour.

General purpose chatbots answer every question fully and fluently, which is exactly what a learner does not need: a finished answer skips the struggle that produces learning, and it makes cheating trivial. The job is different from a help desk. A tutor has to hold back, ask the next question, spot the misconception and push the student to do the work, within the course's rules on academic honesty.

The evidence so far is mixed. A randomized trial of Khan Academy's Khanmigo in 18 Tennessee middle schools found small gains that resembled those from Khan Academy practice without AI, and that almost every student tried the tutor but rarely engaged it in substantive mathematical dialogue; the authors suggest low engagement as one explanation. A World Bank pilot in Nigeria, run with teacher support, reported large gains in six weeks, and larger gains for students who attended more sessions. The Khanmigo authors conclude that realizing the promise of AI tutoring will require getting students to use it, not just giving them access.

How does it work?

  1. Set the scope. The school or teacher defines the course, the material the tutor may use and the rules: no full solutions to graded work, adherence to the academic honesty policy, escalation topics such as wellbeing concerns.
  2. Ground the tutor in the course. Lecture notes, readings and worked examples are indexed so the tutor explains in the course's own terms and cites where a concept is taught.
  3. Coach, do not solve. The tutor asks what the student has tried, gives the smallest useful hint, checks understanding with a question and only then moves on. A second check reviews each reply before the student sees it and can reject or retry replies that give away an answer.
  4. Limit and pace use. A cap on questions per period (CS50 uses a "heart" system) stops students from replacing thinking with hundreds of prompts, and prompts inside the exercise flow nudge students who are stuck but not asking.
  5. Keep the teacher in the loop. Teachers see aggregated topics and misconceptions, can read conversations under the school's policy, and take over for anything personal or sensitive.
Audience
Customer facing
Autonomy
Assist
Adoption
Emerging
Channels
Web chat, Mobile app

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

No public deployment has disclosed a measurable outcome yet.

Value drivers: Inclusion and access, Customer experience, Employee productivity.

Indicative value

A school district with 10,000 students in grades 6 to 12

USD 15,000 to USD 840,000

Equivalent value of tutoring time delivered per year

How this is calculated

Formula: students * activeShare * hoursPerStudent * equivalence * tutorCostPerHour. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Students with access to the tutor students, students10,00010,000The reference district.
Share of students who use the tutor substantively activeShare, fraction of students0.150.35Conservative on purpose. In the Khanmigo trial on this page, 96 percent of students tried the tutor but the median student messaged it on only a third of practice days. Editorial assumption, replace with your own usage data.
Tutoring hours per active student per year hoursPerStudent, hours per student per year515Editorial assumption, replace with your own usage data.
Value of an AI tutoring hour relative to a human tutoring hour equivalence, fraction of a human tutoring hour0.10.4Editorial assumption. The Khanmigo trial on this page found gains similar to Khan Academy practice without AI, so the trial evidence does not support treating an AI tutoring hour as equal to a human one.
Cost of an hour of human tutoring tutorCostPerHour, USD per hour2040Editorial assumption for group or online tutoring, replace with your local rate.

What it leaves out: A proxy, not a learning outcome. It prices tutoring time, discounted heavily because AI tutoring is not equivalent to a human tutor, and leaves out licence and integration costs, teacher time to supervise and the risk that students use the tutor to avoid work. The larger randomized trial on this page (Khanmigo, 18 Tennessee middle schools) found no clear gain from the AI tutor beyond what the same practice platform delivers without it, so the value may be close to zero where students rarely engage; the positive Nigerian pilot was short, ran with teacher support and was reported by the World Bank team ahead of formal publication. Measure learning gains against a comparison group before claiming more.

Who already uses it?

3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Hamilton County Schools

United States · Education · 2024

PilotGrade B

Philip Oreopoulos and Nina Low ran a two year cluster randomized trial in 18 Tennessee middle schools in Hamilton County in the 2024/25 and 2025/26 school years, in which randomly assigned students used Khan Academy with its AI tutor Khanmigo, configured to coach rather than give answers, during existing daily remedial maths sessions. Assignment raised maths achievement by 1.3 national percentile ranks per term, similar to Khan Academy practice without AI. Almost every student tried Khanmigo, but the median student messaged it on only a third of the days they practiced and in only 17 percent of the exercise sessions in which they made a mistake, and the messages students did send were mostly bare answers or clicks on suggested prompts. Chalkbeat reports that Khan Academy has since redesigned its interface to integrate Khanmigo better.

No outcome disclosed.

World Bank

Nigeria · Education · 2024

PilotGrade B

In June and July 2024 a World Bank team, working with the Edo State education authorities, ran a six week after school programme in Edo, Nigeria, in which 800 first year senior secondary students used Microsoft Copilot as a tutor, mainly to learn English, with teacher support: teachers introduced each session's topic, suggested prompts and mentored students as they worked with the tool. In a randomized evaluation, participants outperformed their peers in English, AI knowledge and digital skills, and also did better in their end of year exams. The team reports learning gains of about 0.3 standard deviations, says the programme outperformed 80% of the interventions in a database of randomized evaluations in developing countries, and reports that the more sessions students attended, the greater their gains; girls, who started behind boys, seemed to gain even more.

No outcome disclosed.

Harvard University

United States · Education · 2023

ProductionGrade B

Since spring 2023 Harvard's introductory computer science course CS50 has run the "CS50 Duck", an AI tutor built on the ChatGPT API that is deliberately less helpful than a general chatbot: it asks more questions than it answers and avoids giving outright solutions. The course combines prompting with its own code that tries to evaluate each reply before the student sees it and sometimes rejects or retries it, and added a "heart system" that limits questions per period after some students asked around 200 questions. David Malan calls it a net positive but acknowledges the Duck still sometimes returns code despite instructions not to.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Course materials, worked examples and rubrics the school has the right to use
  • The course's academic honesty policy written as rules the tutor can follow
  • A set of real student questions and misconceptions to test against

Systems to integrate

  • Learning management system or course platform (single sign on, course roster)
  • Exercise or practice platform, so the tutor can appear where students get stuck
  • Safeguarding and wellbeing referral process

Complexity: Medium

A chatbot is easy; a tutor that reliably holds back answers, stays inside the course and is used well by students is hard. Most of the work is pedagogy, safeguarding and embedding the tutor in the exercise flow, not integration.

  1. 1

    Start with one course and one teacher team

    Pick a course with high demand for help and a teacher team willing to shape the tutor's behaviour. Write down what the tutor must never do (give a full solution to graded work) and what it should always do (ask what the student tried).

  2. 2

    Ground it in the course

    Index the course's own notes, readings and examples, and make the tutor cite where a concept is taught. Refuse or redirect questions outside the course.

  3. 3

    Add an answer check

    Review every reply before the student sees it with a second pass that rejects replies that contain full solutions or break the honesty policy. CS50 reports that instructions alone were not enough.

  4. 4

    Put the tutor where students get stuck

    Embed it in the exercise flow and prompt it after a wrong answer, rather than as a separate chat tab students must choose to open. The Khanmigo trial shows that optional access alone produces little use.

  5. 5

    Pace use and involve teachers

    Cap questions per student per period, show teachers the common misconceptions each week and agree how teachers follow up with students who over rely on the tutor or show signs of distress.

  6. 6

    Evaluate learning, not usage

    Compare learning outcomes with a comparison group over at least a term, and report engagement honestly, including off topic use and attempts to extract answers.

Guardrails

  • A reply check that blocks full solutions to graded work and anything against the honesty policy
  • Answers grounded in approved course material, with a refusal outside the course
  • A cap on questions per student per period
  • Age appropriate content filters and escalation of wellbeing or safeguarding signals to staff
  • No emotion recognition of students, which the EU AI Act prohibits in education

KPIs to instrument

  • Share of students who use the tutor substantively each week, not just once
  • Share of stuck moments (wrong answers) in which the student asks the tutor for help
  • Replies blocked or retried by the answer check
  • Learning gains against a comparison group over a term
  • Student and teacher satisfaction

Human in the loop

Teachers own the course scope, the rules and the follow up. They review aggregated topics and a sample of conversations each week, handle any safeguarding signal, and decide grades; the tutor never grades or places a student.

Common failure modes

Access without engagement
Students try the tutor once and stop, or ask it off topic questions. Embed it in the work and prompt it at the moment of error.
Answer vending
Students talk the tutor into giving the solution, which removes the learning. Check replies before they are shown and log extraction attempts.
Over reliance
A few students ask hundreds of questions instead of thinking. Cap questions per period and let teachers follow up.
Confidently wrong explanations
The tutor explains a concept incorrectly. Ground it in course material, test it on known misconceptions and let students flag errors to teachers.

What are the risks and rules?

EU AI Act

Depends on design

A tutor that only converses with students falls under the transparency duty of Article 50. It becomes high risk under Annex III point 3(b) when it evaluates learning outcomes, including when those outcomes are used to steer a student's learning process, and under point 3(c) when it assesses the level of education a student should receive. Inferring students' emotions is prohibited in education institutions under Article 5(1)(f).

Guidance

Controls to put in place

  • AI disclosure to students and parents, with the school's rules for use
  • A data protection impact assessment covering minors and conversation logs
  • Inventory entry with an accountable owner per course
  • Regression tests for answer giving and off topic behaviour on every prompt or model change
  • Teacher review of aggregated topics and a sample of conversations

Frequently asked questions

Do AI tutors improve learning?
Sometimes, and it depends on how they are used. A World Bank pilot in Edo, Nigeria, run with teacher support, reported gains of about 0.3 standard deviations in six weeks. A two year randomized trial of Khanmigo in 18 Tennessee middle schools found small gains similar to Khan Academy practice without AI; the authors suggest low engagement as one explanation, as the median student messaged the tutor on only a third of the days they practiced.
How do you stop an AI tutor from giving students the answers?
Instructions alone are not enough. Harvard's CS50 combines prompting with code that tries to evaluate each reply before the student sees it and sometimes rejects or retries it, plus a limit on questions per period. Test the tutor against real attempts to extract answers on every change.
Is an AI tutor high risk under the EU AI Act?
A tutor that only converses is subject to the transparency duty. It is high risk under Annex III point 3(b) if it evaluates learning outcomes, including when those outcomes steer a student's learning, or under point 3(c) if it assesses the level of education a student should receive. Inferring students' emotions in education is prohibited outright.

How to cite this page

Blits.ai AI Use Case Library, "AI tutor that coaches students through problems", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/ai-tutor-for-students. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Education

AI academic advising assistant for course selection and degree requirements

An AI assistant that answers students' questions about degree requirements, course selection, prerequisites and majors, grounded in the institution's own catalog and advising documents, so students get quick answers to routine questions and are directed to a human advisor for anything that needs judgment, is time sensitive, or falls outside what the assistant can see.

Deployments
3 public, best grade B
Autonomy
Assist
Education

AI assistant for student enrollment and student services

An AI assistant that answers admitted and current students' questions about admissions, financial aid, registration, housing and deadlines by text message and web chat, sends timely reminders for the tasks each student still has to complete, and hands personal or complex cases to staff.

Deployments
3 public, best grade B
Autonomy
Supervised agent
BankingPayments and cards

AI agent for account and card servicing

An AI agent that resolves routine account and card requests end to end, such as balances, statements, card blocks and replacements, PIN resets and limit changes, across app, web, messaging and phone, and hands anything sensitive or unusual to a human with the full context.

Deployments
2 public, best grade B
Reported containment rate
about 90%
DBS Bank, organization claim
Real estate

AI agent for apartment leasing inquiries and resident service

An AI agent that answers rental prospects and residents by chat, text, email and phone for a property manager: it answers questions about apartments and policies, books tours, takes maintenance requests, sends renewal and payment reminders, and hands anything that needs judgment to leasing or service staff.

Deployments
3 public, best grade B
Autonomy
Supervised agent