What problem does it solve?
A 2026 working paper on AI tutoring sums up the research on human tutoring: the best one to one programmes produce gains of a third of a standard deviation or more, but high dosage tutoring often costs several thousand dollars per student per year. In large courses the queue for help is the bottleneck: Harvard's CS50 recalls times when office hours became unmanageable and the average wait could be as long as an hour.
General purpose chatbots answer every question fully and fluently, which is exactly what a learner does not need: a finished answer skips the struggle that produces learning, and it makes cheating trivial. The job is different from a help desk. A tutor has to hold back, ask the next question, spot the misconception and push the student to do the work, within the course's rules on academic honesty.
The evidence so far is mixed. A randomized trial of Khan Academy's Khanmigo in 18 Tennessee middle schools found small gains that resembled those from Khan Academy practice without AI, and that almost every student tried the tutor but rarely engaged it in substantive mathematical dialogue; the authors suggest low engagement as one explanation. A World Bank pilot in Nigeria, run with teacher support, reported large gains in six weeks, and larger gains for students who attended more sessions. The Khanmigo authors conclude that realizing the promise of AI tutoring will require getting students to use it, not just giving them access.
- A 2026 EdWorkingPaper by Philip Oreopoulos and Nina Low, citing Nickow et al. (2024), puts the average effect of tutoring programmes at roughly 0.3 standard deviations, and notes that high dosage programmes often cost several thousand dollars per student per year.One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment (2026)
How does it work?
- Set the scope. The school or teacher defines the course, the material the tutor may use and the rules: no full solutions to graded work, adherence to the academic honesty policy, escalation topics such as wellbeing concerns.
- Ground the tutor in the course. Lecture notes, readings and worked examples are indexed so the tutor explains in the course's own terms and cites where a concept is taught.
- Coach, do not solve. The tutor asks what the student has tried, gives the smallest useful hint, checks understanding with a question and only then moves on. A second check reviews each reply before the student sees it and can reject or retry replies that give away an answer.
- Limit and pace use. A cap on questions per period (CS50 uses a "heart" system) stops students from replacing thinking with hundreds of prompts, and prompts inside the exercise flow nudge students who are stuck but not asking.
- Keep the teacher in the loop. Teachers see aggregated topics and misconceptions, can read conversations under the school's policy, and take over for anything personal or sensitive.
- Audience
- Customer facing
- Autonomy
- Assist
- Adoption
- Emerging
- Channels
- Web chat, Mobile app
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
No public deployment has disclosed a measurable outcome yet.
Value drivers: Inclusion and access, Customer experience, Employee productivity.
Indicative value
A school district with 10,000 students in grades 6 to 12
USD 15,000 to USD 840,000
Equivalent value of tutoring time delivered per year
How this is calculated
Formula: students * activeShare * hoursPerStudent * equivalence * tutorCostPerHour. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Students with access to the tutor students, students | 10,000 | 10,000 | The reference district. |
| Share of students who use the tutor substantively activeShare, fraction of students | 0.15 | 0.35 | Conservative on purpose. In the Khanmigo trial on this page, 96 percent of students tried the tutor but the median student messaged it on only a third of practice days. Editorial assumption, replace with your own usage data. |
| Tutoring hours per active student per year hoursPerStudent, hours per student per year | 5 | 15 | Editorial assumption, replace with your own usage data. |
| Value of an AI tutoring hour relative to a human tutoring hour equivalence, fraction of a human tutoring hour | 0.1 | 0.4 | Editorial assumption. The Khanmigo trial on this page found gains similar to Khan Academy practice without AI, so the trial evidence does not support treating an AI tutoring hour as equal to a human one. |
| Cost of an hour of human tutoring tutorCostPerHour, USD per hour | 20 | 40 | Editorial assumption for group or online tutoring, replace with your local rate. |
What it leaves out: A proxy, not a learning outcome. It prices tutoring time, discounted heavily because AI tutoring is not equivalent to a human tutor, and leaves out licence and integration costs, teacher time to supervise and the risk that students use the tutor to avoid work. The larger randomized trial on this page (Khanmigo, 18 Tennessee middle schools) found no clear gain from the AI tutor beyond what the same practice platform delivers without it, so the value may be close to zero where students rarely engage; the positive Nigerian pilot was short, ran with teacher support and was reported by the World Bank team ahead of formal publication. Measure learning gains against a comparison group before claiming more.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Hamilton County Schools
United States · Education · 2024
Philip Oreopoulos and Nina Low ran a two year cluster randomized trial in 18 Tennessee middle schools in Hamilton County in the 2024/25 and 2025/26 school years, in which randomly assigned students used Khan Academy with its AI tutor Khanmigo, configured to coach rather than give answers, during existing daily remedial maths sessions. Assignment raised maths achievement by 1.3 national percentile ranks per term, similar to Khan Academy practice without AI. Almost every student tried Khanmigo, but the median student messaged it on only a third of the days they practiced and in only 17 percent of the exercise sessions in which they made a mistake, and the messages students did send were mostly bare answers or clicks on suggested prompts. Chalkbeat reports that Khan Academy has since redesigned its interface to integrate Khanmigo better.
No outcome disclosed.
World Bank
Nigeria · Education · 2024
In June and July 2024 a World Bank team, working with the Edo State education authorities, ran a six week after school programme in Edo, Nigeria, in which 800 first year senior secondary students used Microsoft Copilot as a tutor, mainly to learn English, with teacher support: teachers introduced each session's topic, suggested prompts and mentored students as they worked with the tool. In a randomized evaluation, participants outperformed their peers in English, AI knowledge and digital skills, and also did better in their end of year exams. The team reports learning gains of about 0.3 standard deviations, says the programme outperformed 80% of the interventions in a database of randomized evaluations in developing countries, and reports that the more sessions students attended, the greater their gains; girls, who started behind boys, seemed to gain even more.
No outcome disclosed.
Harvard University
United States · Education · 2023
Since spring 2023 Harvard's introductory computer science course CS50 has run the "CS50 Duck", an AI tutor built on the ChatGPT API that is deliberately less helpful than a general chatbot: it asks more questions than it answers and avoids giving outright solutions. The course combines prompting with its own code that tries to evaluate each reply before the student sees it and sometimes rejects or retries it, and added a "heart system" that limits questions per period after some students asked around 200 questions. David Malan calls it a net positive but acknowledges the Duck still sometimes returns code despite instructions not to.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Course materials, worked examples and rubrics the school has the right to use
- The course's academic honesty policy written as rules the tutor can follow
- A set of real student questions and misconceptions to test against
Systems to integrate
- Learning management system or course platform (single sign on, course roster)
- Exercise or practice platform, so the tutor can appear where students get stuck
- Safeguarding and wellbeing referral process
Complexity: Medium
A chatbot is easy; a tutor that reliably holds back answers, stays inside the course and is used well by students is hard. Most of the work is pedagogy, safeguarding and embedding the tutor in the exercise flow, not integration.
- 1
Start with one course and one teacher team
Pick a course with high demand for help and a teacher team willing to shape the tutor's behaviour. Write down what the tutor must never do (give a full solution to graded work) and what it should always do (ask what the student tried).
- 2
Ground it in the course
Index the course's own notes, readings and examples, and make the tutor cite where a concept is taught. Refuse or redirect questions outside the course.
- 3
Add an answer check
Review every reply before the student sees it with a second pass that rejects replies that contain full solutions or break the honesty policy. CS50 reports that instructions alone were not enough.
- 4
Put the tutor where students get stuck
Embed it in the exercise flow and prompt it after a wrong answer, rather than as a separate chat tab students must choose to open. The Khanmigo trial shows that optional access alone produces little use.
- 5
Pace use and involve teachers
Cap questions per student per period, show teachers the common misconceptions each week and agree how teachers follow up with students who over rely on the tutor or show signs of distress.
- 6
Evaluate learning, not usage
Compare learning outcomes with a comparison group over at least a term, and report engagement honestly, including off topic use and attempts to extract answers.
Guardrails
- A reply check that blocks full solutions to graded work and anything against the honesty policy
- Answers grounded in approved course material, with a refusal outside the course
- A cap on questions per student per period
- Age appropriate content filters and escalation of wellbeing or safeguarding signals to staff
- No emotion recognition of students, which the EU AI Act prohibits in education
KPIs to instrument
- Share of students who use the tutor substantively each week, not just once
- Share of stuck moments (wrong answers) in which the student asks the tutor for help
- Replies blocked or retried by the answer check
- Learning gains against a comparison group over a term
- Student and teacher satisfaction
Human in the loop
Teachers own the course scope, the rules and the follow up. They review aggregated topics and a sample of conversations each week, handle any safeguarding signal, and decide grades; the tutor never grades or places a student.
Common failure modes
- Access without engagement
- Students try the tutor once and stop, or ask it off topic questions. Embed it in the work and prompt it at the moment of error.
- Answer vending
- Students talk the tutor into giving the solution, which removes the learning. Check replies before they are shown and log extraction attempts.
- Over reliance
- A few students ask hundreds of questions instead of thinking. Cap questions per period and let teachers follow up.
- Confidently wrong explanations
- The tutor explains a concept incorrectly. Ground it in course material, test it on known misconceptions and let students flag errors to teachers.
What are the risks and rules?
EU AI Act
Depends on design
A tutor that only converses with students falls under the transparency duty of Article 50. It becomes high risk under Annex III point 3(b) when it evaluates learning outcomes, including when those outcomes are used to steer a student's learning process, and under point 3(c) when it assesses the level of education a student should receive. Inferring students' emotions is prohibited in education institutions under Article 5(1)(f).
Guidance
- Guidance for generative AI in education and research (UNESCO, Global). UNESCO's first global guidance on generative AI in education; it proposes protecting learners' data privacy and setting an age limit for independent conversations with generative AI platforms.
- Generative artificial intelligence (AI) in education (UK Department for Education, Europe). The department's position on generative AI tools in schools and colleges, to be read together with its product safety expectations for generative AI.
- Regulation (EU) 2024/1689 (AI Act), Annex III point 3, education and vocational training (European Union, Europe). Lists as high risk AI that evaluates learning outcomes, including when those outcomes steer the learning process, and AI that assesses the appropriate level of education a person will receive (point 3(b) and 3(c)).
- Article 5, prohibited AI practices (European Union, Europe). Prohibits AI systems that infer the emotions of a natural person in education institutions, except for medical or safety reasons.
Controls to put in place
- AI disclosure to students and parents, with the school's rules for use
- A data protection impact assessment covering minors and conversation logs
- Inventory entry with an accountable owner per course
- Regression tests for answer giving and off topic behaviour on every prompt or model change
- Teacher review of aggregated topics and a sample of conversations
Frequently asked questions
- Do AI tutors improve learning?
- Sometimes, and it depends on how they are used. A World Bank pilot in Edo, Nigeria, run with teacher support, reported gains of about 0.3 standard deviations in six weeks. A two year randomized trial of Khanmigo in 18 Tennessee middle schools found small gains similar to Khan Academy practice without AI; the authors suggest low engagement as one explanation, as the median student messaged the tutor on only a third of the days they practiced.
- How do you stop an AI tutor from giving students the answers?
- Instructions alone are not enough. Harvard's CS50 combines prompting with code that tries to evaluate each reply before the student sees it and sometimes rejects or retries it, plus a limit on questions per period. Test the tutor against real attempts to extract answers on every change.
- Is an AI tutor high risk under the EU AI Act?
- A tutor that only converses is subject to the transparency duty. It is high risk under Annex III point 3(b) if it evaluates learning outcomes, including when those outcomes steer a student's learning, or under point 3(c) if it assesses the level of education a student should receive. Inferring students' emotions in education is prohibited outright.
How to cite this page
Blits.ai AI Use Case Library, "AI tutor that coaches students through problems", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/ai-tutor-for-students. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published