AI use case

AI assisted proctoring and integrity monitoring for remote exams

AI that supports the integrity of a remote, high stakes exam by verifying a test taker's identity, analysing behaviour such as typing patterns, facial matching and session activity for signs of impersonation or unauthorized help, and flagging sessions for a trained human reviewer to decide, rather than letting a model issue an automated finding of misconduct on its own.

By Len Debets · Last verified 29 September 2026 · 2 public deployments

USD 800,000 to USD 4 million
Indicative value per year
A testing organization administering 500,000 remote exam sessions a year. Worked example, see how it is calculated.

What problem does it solve?

High stakes tests, from language proficiency exams to graduate admission tests and teacher licensing exams, such as the GRE, TOEFL and Praxis, increasingly happen at home rather than in a supervised test centre, a shift ETS says the Covid 19 pandemic forced within weeks. Remote testing removes the travel and scheduling burden of a test centre, but it also removes the one thing a physical test centre offered by default: a proctor who can see the whole room.

A single human proctor watching a live video feed cannot reliably catch a hidden phone, a second screen, a coached answer or a subtle handoff, and cannot review what happened after the fact unless the session was recorded. At the same time, evidence has piled up that automated cheating detection is not reliable enough to stand alone: several institutions, including Yale, Vanderbilt and Johns Hopkins Universities, disabled Turnitin's AI writing detection feature, citing concerns about false positives. The honest state of the art pairs AI generated signals with a trained human decision, and treats an AI flag as a lead to investigate, not a verdict.

How does it work?

  1. Verify identity before the exam starts. Government ID checks, facial matching and, on some exams, voice biometrics confirm the person taking the test is who they registered as, and repeat through the session to catch a substitution.
  2. Analyse behaviour throughout the session. AI systems track keystroke and typing rhythm, keep the same person confirmed in the seat and watch for a second person entering the room, and monitor application or browser activity for patterns associated with outside help or content that was memorized or copied rather than produced live.
  3. Look across sessions, not just one. Some systems, such as the Duolingo English Test, analyse patterns across thousands of test sessions at once, for example near identical answers from test takers in the same region, which a single proctor watching one session could never see.
  4. Flag, do not decide. The AI signals are handed to a trained human reviewer as a diagnosis to investigate, with the full session recording available to pause, rewind and examine in context, rather than triggering an automatic finding of misconduct.
  5. Record for later review. The full session is recorded, so a flagged case, a dispute or an audit can be reviewed after the fact, which a live only human proctor in a test centre cannot offer.
Audience
Back office
Autonomy
Supervised agent
Adoption
Mainstream
Channels
Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

No public deployment has disclosed a measurable outcome yet.

Value drivers: Risk and loss reduction, Compliance quality, Inclusion and access.

Indicative value

A testing organization administering 500,000 remote exam sessions a year

USD 800,000 to USD 4 million

Proctoring cost avoided versus one to one live human proctoring per year

How this is calculated

Formula: sessions * costPerHumanOnlyProctor * costReductionShare. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Remote exam sessions per year sessions, sessions per year500,000500,000The reference organization.
Cost of a fully live, one proctor per session model costPerHumanOnlyProctor, USD per session820Editorial assumption for a live, one to one remote human proctor; replace with your own proctoring contract rates.
Share of proctoring cost avoided by asynchronous, AI assisted review instead of live one to one proctoring costReductionShare, fraction of cost0.20.4Editorial assumption. No source in this page's evidence discloses a cost figure; ETS and Duolingo describe the security approach, not its cost, so this range is a hypothesis to test against your own proctoring contract, not a benchmark.

What it leaves out: Entirely a modelled hypothesis with no disclosed cost benchmark behind it. It leaves out the cost of the AI systems themselves, the reviewer time still needed for every flagged session, and the cost of a false accusation or a missed cheating case, which for a high stakes exam can be far larger than the proctoring cost itself.

Who already uses it?

2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Duolingo

United States · Education · 2026

ProductionGrade B

The Duolingo English Test, a remote, at home English proficiency test, secures every session with a proprietary lockdown testing application, layered identity verification (government ID, facial matching, device fingerprinting), and AI assisted behaviour analysis, including keystroke and typing rhythm monitoring and cross session pattern detection. Duolingo states that AI flags suspicious sessions for review but that "decisions aren't automated": a trained human proctor individually reviews every test session, with the ability to pause, speed up or slow down the recording, before any action follows.

No outcome disclosed.

Educational Testing Service (ETS)

United States · Education · 2020

ProductionGrade B

ETS built a remote proctored testing option for the GRE General Test and TOEFL iBT test in six weeks in 2020, in collaboration with ProctorU, later extending it to the HiSET exam and Praxis tests. The at home option layers identity verification (ID matching and, for TOEFL, voice biometrics), a live remote proctor who conducts a 360 degree room scan and monitors the session, and artificial intelligence that continuously scans for irregularities such as a second person entering the room, working alongside process data analysis by ETS's research and psychometrics team to detect suspicious patterns over time.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • A defined set of behaviours that count as a security signal, agreed with academic or legal stakeholders
  • Recorded video, audio and process data for every session, held under a clear retention policy
  • A calibration set of known good and known problematic sessions to test human reviewer judgment
  • Documented accommodations for test takers who cannot meet standard camera, room or biometric requirements

Systems to integrate

  • Identity verification and biometric matching provider
  • Secure testing application or lockdown browser
  • Session recording and case management for flagged reviews
  • Test delivery platform and score reporting system

Complexity: High

Identity verification and basic recording are straightforward to buy from a proctoring vendor. The hard part is everything the flag touches afterward: a human review process trained and calibrated so it does not simply rubber stamp or blindly reject AI signals, a defensible appeals process, and a bias and accessibility review of biometric checks, since accuracy is not guaranteed to be even across skin tones or for test takers with some disabilities and should be tested on your own population rather than assumed.

  1. 1

    Decide what a flag means before you launch

    Write down, before go live, exactly what a flagged session leads to: a human review, a request for clarification, a retest, or a finding, and who can override the AI's signal at each step.

  2. 2

    Calibrate human reviewers against automation bias

    A reviewer can lean on an AI flag instead of independently checking the recording. Require guidance that asks for independent evidence in the recording, not the AI signal alone, and test reviewers against known false signals before launch and periodically after. Do not assume a human in the loop fixes accuracy on its own.

  3. 3

    Keep the full recording, not just the flag

    A flag without the underlying video or process data cannot be investigated or defended in a dispute. Store the full session so a reviewer, and if needed an appeals panel, can see the context.

  4. 4

    Separate detection from adjudication

    The team or system that raises a flag should not be the same one that issues a finding of misconduct against a student, mirroring how a fraud alert and a fraud decision are usually kept apart in other industries.

  5. 5

    Publish the appeal path and use it

    Test takers need a real, timely way to contest a flag or a finding, with a person who was not involved in the original flag reviewing it.

Guardrails

  • No automated finding of misconduct without a trained human reviewer examining the recording
  • A documented, tested appeal path, reviewed by someone not involved in the original flag
  • Accommodations for test takers who cannot meet standard camera, room or biometric requirements
  • Retention and access limits on biometric data, video and process data of identifiable people

KPIs to instrument

  • False positive and false negative rate of the AI signal against reviewer decisions, on a calibration set
  • Reviewer agreement rate and how often reviewers accept an AI signal without independent evidence
  • Time from flag to human decision
  • Appeals received, upheld and their outcomes

Human in the loop

A trained human reviewer examines every flagged session before a finding of misconduct is recorded, and can pause, rewind and review the recording rather than relying on a real time glance alone. ETS describes a live proctor watching every session in real time, with the ability to cancel a session immediately if there is a clear attempt to cheat, pending review, while Duolingo describes a trained proctor reviewing every session's recording anonymously and asynchronously afterward. Both treat AI as an aid to that person's judgment, not a replacement for it.

Common failure modes

Automation bias in the human reviewer
A reviewer trusts the AI's flag instead of independently checking the recording, accepting it more often than an accurate signal alone would justify. Test and monitor reviewer behaviour against known false signals, not only model accuracy.
Detection tool disabled after real world false positives
Several universities, including Yale, Vanderbilt and Johns Hopkins, turned off Turnitin's AI writing detector after it produced false positives that undermined confidence in it; a tool that is not trusted stops being used. Validate detection accuracy on your own population before relying on it for a consequence.
Biometric and behavioural checks that fail some test takers more than others
Facial matching and behavioural analysis are not guaranteed to perform evenly across skin tones, ages or for test takers with certain disabilities, which can produce more false flags for those groups if left untested. Test accuracy across demographic groups on your own population and provide a fallback for accommodations.
A false accusation that is never corrected
An incorrect finding of misconduct can end a student's exam, admission or licensure prospects. A slow or absent appeal path turns a detection error into a life changing one. Time bound the appeal process and staff it independently of the original decision.

What are the risks and rules?

EU AI Act

High risk

Annex III point 3(d) lists AI systems intended to be used for monitoring and detecting prohibited behaviour of students during tests in the context of education and training institutions as high risk. One to one biometric identity verification, confirming a test taker is who they claim to be, falls outside the biometric identification use covered by Annex III point 1(a). The Article 5 restriction that does apply is different: Article 5(1)(f) prohibits AI systems that infer emotions in an education institution, so the behavioural analysis in this use case must stay limited to security signals such as identity and typing patterns, and must not be built or read as emotion inference.

Rules that apply

Guidance

Controls to put in place

  • Human review of the full session recording before any finding of misconduct
  • A timed, independently staffed appeal process
  • Documented accommodations for camera, room and biometric requirements
  • Retention limits and access controls on video, audio and biometric data
  • Accuracy testing of identity and behaviour signals across demographic groups

When it went wrong elsewhere

  • Universities disable Turnitin's AI writing detector over false positives. Several institutions, including Yale, Vanderbilt and Johns Hopkins Universities, have disabled Turnitin's AI detection feature, citing concerns about false positives, The Chronicle of Higher Education reported. The live article returns an HTTP 403 for automated access; this is the Wayback Machine copy, checked 2026-09-29.

Frequently asked questions

Can AI alone determine that a student cheated?
In the systems in this page's evidence, no. Duolingo and ETS both describe a trained human proctor reviewing every session, with AI as an aid to detect signals a person watching in real time would miss, not as the decision maker. Duolingo says in a blog post that decisions "aren't automated."
Why are universities turning off AI cheating detectors?
Because the accuracy has not matched the stakes. Several institutions, including Yale, Vanderbilt and Johns Hopkins, disabled Turnitin's AI writing detection tool after it produced false positives. This is different from remote proctoring's identity and behaviour signals, which are reviewed by a human before a finding is recorded, but the lesson, validate before you rely on it, applies to both.
How does remote AI proctoring compare with a human proctor in a test centre?
ETS says it built its first remote proctored option, with ProctorU, in six weeks in 2020, and has since run it for hundreds of thousands of test takers on the GRE, TOEFL iBT, HiSET and Praxis exams, layering identity checks, environmental scanning and process data analysis on top of live human proctoring. Duolingo argues this catches some cheating methods, such as collusion patterns across thousands of sessions, that a single test centre proctor never could.
Is exam proctoring AI regulated?
Under the EU AI Act, Annex III point 3(d) lists AI used to monitor and detect prohibited student behaviour during tests as high risk, which brings obligations for risk management, human oversight, logging and conformity assessment. Facial recognition used for identity checks can raise separate biometric data questions under GDPR and similar laws.

How to cite this page

Blits.ai AI Use Case Library, "AI assisted proctoring and integrity monitoring for remote exams", last verified 29 September 2026, https://www.blits.ai/ai-use-cases/exam-and-assessment-integrity. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 29 September 2026: First published

Related use cases

EducationGovernment and public sector

AI scoring of essays and written answers in assessments

AI that scores students' essays and short written answers against a rubric, trained on responses scored by human raters, with human raters rescoring a sample of responses and every response the engine is unsure about. In hybrid programmes such as Texas, a human score is the score of record whenever a human scores a response.

Deployments
3 public, best grade B
Autonomy
Supervised agent
Retail and ecommerceEducation

AI video analytics and screening for physical security

AI that watches camera feeds or walk through sensors at a site, flags a likely weapon, intrusion or theft in real time or on search, and leaves the verification and every response action to a human guard or investigator, rather than acting on its own.

Deployments
3 public, best grade C
Autonomy
Assist
Healthcare

AI prioritization of radiology and imaging worklists

An AI system that analyzes a medical image immediately after a scan, flags time sensitive findings such as a brain bleed, a stroke causing large vessel occlusion or a pulmonary embolism, and reorders the radiologist's worklist and notifies the care team so the most urgent cases are read and acted on first, while a radiologist confirms every finding before it changes a patient's treatment.

Deployments
2 public, best grade B
Reported cycle time reduction
about 44%
Adventist Health + Rideout, vendor claim
Education

AI assistant for student enrollment and student services

An AI assistant that answers admitted and current students' questions about admissions, financial aid, registration, housing and deadlines by text message and web chat, sends timely reminders for the tasks each student still has to complete, and hands personal or complex cases to staff.

Deployments
3 public, best grade B
Autonomy
Supervised agent