What problem does it solve?
High stakes tests, from language proficiency exams to graduate admission tests and teacher licensing exams, such as the GRE, TOEFL and Praxis, increasingly happen at home rather than in a supervised test centre, a shift ETS says the Covid 19 pandemic forced within weeks. Remote testing removes the travel and scheduling burden of a test centre, but it also removes the one thing a physical test centre offered by default: a proctor who can see the whole room.
A single human proctor watching a live video feed cannot reliably catch a hidden phone, a second screen, a coached answer or a subtle handoff, and cannot review what happened after the fact unless the session was recorded. At the same time, evidence has piled up that automated cheating detection is not reliable enough to stand alone: several institutions, including Yale, Vanderbilt and Johns Hopkins Universities, disabled Turnitin's AI writing detection feature, citing concerns about false positives. The honest state of the art pairs AI generated signals with a trained human decision, and treats an AI flag as a lead to investigate, not a verdict.
How does it work?
- Verify identity before the exam starts. Government ID checks, facial matching and, on some exams, voice biometrics confirm the person taking the test is who they registered as, and repeat through the session to catch a substitution.
- Analyse behaviour throughout the session. AI systems track keystroke and typing rhythm, keep the same person confirmed in the seat and watch for a second person entering the room, and monitor application or browser activity for patterns associated with outside help or content that was memorized or copied rather than produced live.
- Look across sessions, not just one. Some systems, such as the Duolingo English Test, analyse patterns across thousands of test sessions at once, for example near identical answers from test takers in the same region, which a single proctor watching one session could never see.
- Flag, do not decide. The AI signals are handed to a trained human reviewer as a diagnosis to investigate, with the full session recording available to pause, rewind and examine in context, rather than triggering an automatic finding of misconduct.
- Record for later review. The full session is recorded, so a flagged case, a dispute or an audit can be reviewed after the fact, which a live only human proctor in a test centre cannot offer.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Mainstream
- Channels
- Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
No public deployment has disclosed a measurable outcome yet.
Value drivers: Risk and loss reduction, Compliance quality, Inclusion and access.
Indicative value
A testing organization administering 500,000 remote exam sessions a year
USD 800,000 to USD 4 million
Proctoring cost avoided versus one to one live human proctoring per year
How this is calculated
Formula: sessions * costPerHumanOnlyProctor * costReductionShare. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Remote exam sessions per year sessions, sessions per year | 500,000 | 500,000 | The reference organization. |
| Cost of a fully live, one proctor per session model costPerHumanOnlyProctor, USD per session | 8 | 20 | Editorial assumption for a live, one to one remote human proctor; replace with your own proctoring contract rates. |
| Share of proctoring cost avoided by asynchronous, AI assisted review instead of live one to one proctoring costReductionShare, fraction of cost | 0.2 | 0.4 | Editorial assumption. No source in this page's evidence discloses a cost figure; ETS and Duolingo describe the security approach, not its cost, so this range is a hypothesis to test against your own proctoring contract, not a benchmark. |
What it leaves out: Entirely a modelled hypothesis with no disclosed cost benchmark behind it. It leaves out the cost of the AI systems themselves, the reviewer time still needed for every flagged session, and the cost of a false accusation or a missed cheating case, which for a high stakes exam can be far larger than the proctoring cost itself.
Who already uses it?
2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Duolingo
United States · Education · 2026
The Duolingo English Test, a remote, at home English proficiency test, secures every session with a proprietary lockdown testing application, layered identity verification (government ID, facial matching, device fingerprinting), and AI assisted behaviour analysis, including keystroke and typing rhythm monitoring and cross session pattern detection. Duolingo states that AI flags suspicious sessions for review but that "decisions aren't automated": a trained human proctor individually reviews every test session, with the ability to pause, speed up or slow down the recording, before any action follows.
No outcome disclosed.
Educational Testing Service (ETS)
United States · Education · 2020
ETS built a remote proctored testing option for the GRE General Test and TOEFL iBT test in six weeks in 2020, in collaboration with ProctorU, later extending it to the HiSET exam and Praxis tests. The at home option layers identity verification (ID matching and, for TOEFL, voice biometrics), a live remote proctor who conducts a 360 degree room scan and monitors the session, and artificial intelligence that continuously scans for irregularities such as a second person entering the room, working alongside process data analysis by ETS's research and psychometrics team to detect suspicious patterns over time.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- A defined set of behaviours that count as a security signal, agreed with academic or legal stakeholders
- Recorded video, audio and process data for every session, held under a clear retention policy
- A calibration set of known good and known problematic sessions to test human reviewer judgment
- Documented accommodations for test takers who cannot meet standard camera, room or biometric requirements
Systems to integrate
- Identity verification and biometric matching provider
- Secure testing application or lockdown browser
- Session recording and case management for flagged reviews
- Test delivery platform and score reporting system
Complexity: High
Identity verification and basic recording are straightforward to buy from a proctoring vendor. The hard part is everything the flag touches afterward: a human review process trained and calibrated so it does not simply rubber stamp or blindly reject AI signals, a defensible appeals process, and a bias and accessibility review of biometric checks, since accuracy is not guaranteed to be even across skin tones or for test takers with some disabilities and should be tested on your own population rather than assumed.
- 1
Decide what a flag means before you launch
Write down, before go live, exactly what a flagged session leads to: a human review, a request for clarification, a retest, or a finding, and who can override the AI's signal at each step.
- 2
Calibrate human reviewers against automation bias
A reviewer can lean on an AI flag instead of independently checking the recording. Require guidance that asks for independent evidence in the recording, not the AI signal alone, and test reviewers against known false signals before launch and periodically after. Do not assume a human in the loop fixes accuracy on its own.
- 3
Keep the full recording, not just the flag
A flag without the underlying video or process data cannot be investigated or defended in a dispute. Store the full session so a reviewer, and if needed an appeals panel, can see the context.
- 4
Separate detection from adjudication
The team or system that raises a flag should not be the same one that issues a finding of misconduct against a student, mirroring how a fraud alert and a fraud decision are usually kept apart in other industries.
- 5
Publish the appeal path and use it
Test takers need a real, timely way to contest a flag or a finding, with a person who was not involved in the original flag reviewing it.
Guardrails
- No automated finding of misconduct without a trained human reviewer examining the recording
- A documented, tested appeal path, reviewed by someone not involved in the original flag
- Accommodations for test takers who cannot meet standard camera, room or biometric requirements
- Retention and access limits on biometric data, video and process data of identifiable people
KPIs to instrument
- False positive and false negative rate of the AI signal against reviewer decisions, on a calibration set
- Reviewer agreement rate and how often reviewers accept an AI signal without independent evidence
- Time from flag to human decision
- Appeals received, upheld and their outcomes
Human in the loop
A trained human reviewer examines every flagged session before a finding of misconduct is recorded, and can pause, rewind and review the recording rather than relying on a real time glance alone. ETS describes a live proctor watching every session in real time, with the ability to cancel a session immediately if there is a clear attempt to cheat, pending review, while Duolingo describes a trained proctor reviewing every session's recording anonymously and asynchronously afterward. Both treat AI as an aid to that person's judgment, not a replacement for it.
Common failure modes
- Automation bias in the human reviewer
- A reviewer trusts the AI's flag instead of independently checking the recording, accepting it more often than an accurate signal alone would justify. Test and monitor reviewer behaviour against known false signals, not only model accuracy.
- Detection tool disabled after real world false positives
- Several universities, including Yale, Vanderbilt and Johns Hopkins, turned off Turnitin's AI writing detector after it produced false positives that undermined confidence in it; a tool that is not trusted stops being used. Validate detection accuracy on your own population before relying on it for a consequence.
- Biometric and behavioural checks that fail some test takers more than others
- Facial matching and behavioural analysis are not guaranteed to perform evenly across skin tones, ages or for test takers with certain disabilities, which can produce more false flags for those groups if left untested. Test accuracy across demographic groups on your own population and provide a fallback for accommodations.
- A false accusation that is never corrected
- An incorrect finding of misconduct can end a student's exam, admission or licensure prospects. A slow or absent appeal path turns a detection error into a life changing one. Time bound the appeal process and staff it independently of the original decision.
What are the risks and rules?
EU AI Act
High risk
Annex III point 3(d) lists AI systems intended to be used for monitoring and detecting prohibited behaviour of students during tests in the context of education and training institutions as high risk. One to one biometric identity verification, confirming a test taker is who they claim to be, falls outside the biometric identification use covered by Annex III point 1(a). The Article 5 restriction that does apply is different: Article 5(1)(f) prohibits AI systems that infer emotions in an education institution, so the behavioural analysis in this use case must stay limited to security signals such as identity and typing patterns, and must not be built or read as emotion inference.
Guidance
- Annex III, high risk AI systems (point 3, education and vocational training) (European Union, Europe). Lists monitoring and detecting prohibited student behaviour during tests as a high risk education use.
- Protecting Student Privacy (US Department of Education, Student Privacy Policy Office, North America). Guidance on FERPA, relevant to video, audio and biometric data collected during a proctored exam.
Controls to put in place
- Human review of the full session recording before any finding of misconduct
- A timed, independently staffed appeal process
- Documented accommodations for camera, room and biometric requirements
- Retention limits and access controls on video, audio and biometric data
- Accuracy testing of identity and behaviour signals across demographic groups
When it went wrong elsewhere
- Universities disable Turnitin's AI writing detector over false positives. Several institutions, including Yale, Vanderbilt and Johns Hopkins Universities, have disabled Turnitin's AI detection feature, citing concerns about false positives, The Chronicle of Higher Education reported. The live article returns an HTTP 403 for automated access; this is the Wayback Machine copy, checked 2026-09-29.
Frequently asked questions
- Can AI alone determine that a student cheated?
- In the systems in this page's evidence, no. Duolingo and ETS both describe a trained human proctor reviewing every session, with AI as an aid to detect signals a person watching in real time would miss, not as the decision maker. Duolingo says in a blog post that decisions "aren't automated."
- Why are universities turning off AI cheating detectors?
- Because the accuracy has not matched the stakes. Several institutions, including Yale, Vanderbilt and Johns Hopkins, disabled Turnitin's AI writing detection tool after it produced false positives. This is different from remote proctoring's identity and behaviour signals, which are reviewed by a human before a finding is recorded, but the lesson, validate before you rely on it, applies to both.
- How does remote AI proctoring compare with a human proctor in a test centre?
- ETS says it built its first remote proctored option, with ProctorU, in six weeks in 2020, and has since run it for hundreds of thousands of test takers on the GRE, TOEFL iBT, HiSET and Praxis exams, layering identity checks, environmental scanning and process data analysis on top of live human proctoring. Duolingo argues this catches some cheating methods, such as collusion patterns across thousands of sessions, that a single test centre proctor never could.
- Is exam proctoring AI regulated?
- Under the EU AI Act, Annex III point 3(d) lists AI used to monitor and detect prohibited student behaviour during tests as high risk, which brings obligations for risk management, human oversight, logging and conformity assessment. Facial recognition used for identity checks can raise separate biometric data questions under GDPR and similar laws.
How to cite this page
Blits.ai AI Use Case Library, "AI assisted proctoring and integrity monitoring for remote exams", last verified 29 September 2026, https://www.blits.ai/ai-use-cases/exam-and-assessment-integrity. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 29 September 2026: First published