What problem does it solve?
Written answers show what students can do in ways multiple choice cannot, but every one has to be read against a rubric by a trained rater. When Texas redesigned its STAAR tests to include more writing, the number of constructed responses to score each year grew six to sevenfold, and the agency estimated that scoring them all by hand would cost 15 to 20 million US dollars more per year. Human scoring is not perfectly consistent either: in a TEA study of extended essays, two trained raters gave exactly the same conventions score on only 67 to 72 percent of responses. In classrooms, the UK Department for Education describes feedback and marking as a burden on teachers.
Automated essay scoring is not new: ETS has used its engine alongside human raters on the GRE since at least 2012, and TEA said in 2024 that at least 21 states use automated scoring for their state assessments. What has changed is scale and scope. State assessment programmes now let an engine give the first score for most written answers, and the UK Department for Education funds AI tools meant to reduce the burden of feedback and marking on teachers. The stakes are high: a wrong score can affect a student's record, a school's rating and public trust, as the Massachusetts rescoring of about 1,400 essays in 2025 showed.
- The Texas Education Agency says the STAAR redesign brought 6 to 7 times more constructed responses to grade each year, and that maintaining full human scoring would have cost 15 to 20 million US dollars more per year.Hybrid Scoring Key Questions (Texas Education Agency, March 2024) (2024)
- In a Texas Education Agency study of spring 2023 STAAR extended constructed responses, two human raters gave exactly the same conventions score on 67 to 72 percent of responses, depending on the item.Hybrid Scoring Key Questions (Texas Education Agency, March 2024) (2024)
How does it work?
- Set the standard with humans. Educators score a set of field test responses against the rubric and agree anchor responses for every score point.
- Train and qualify the engine. The engine is trained on human scored responses for each question and must agree with human raters at the same rate human raters agree with one another, with a similar score distribution, before it is used live.
- Score and flag. The engine gives each response a first score and a confidence value, and flags responses that are blank, too short, off topic, copied, in another language or unlike anything in its training data.
- Route to humans. Flagged and low confidence responses, and a fixed share of all responses, go to trained human raters. Where a human scores a response, the human score is the score of record.
- Monitor and correct. Agreement between engine and humans is monitored daily, preliminary results go to schools with a window to report discrepancies, and rescoring is available.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Mainstream
- Channels
- API and system to system, Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
No public deployment has disclosed a measurable outcome yet.
Value drivers: Lower cost to serve, Speed and cycle time, Employee productivity, Compliance quality.
Indicative value
A state assessment programme scoring 1 million written responses a year
USD 360,000 to USD 1.3 million
Human scoring cost avoided per year
How this is calculated
Formula: responses * engineOnlyShare * humanScoringCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Written responses scored per year responses, responses per year | 1,000,000 | 1,000,000 | The reference programme. |
| Share of responses scored by the engine without a human read engineOnlyShare, fraction of responses | 0.6 | 0.75 | The Texas Education Agency routes at least 25 percent of responses, plus flagged and low confidence ones, to human raters, so at most 75 percent are engine only. The low value is an editorial assumption. |
| Cost of human scoring per response humanScoringCost, USD per response | 0.6 | 1.7 | Derived from the Texas Education Agency deck cited above: 15 to 20 million US dollars a year avoided, spread over the roughly 75 percent of 15.8 million annual responses that the engine now scores alone and that were previously scored by two humans, is about 1.3 to 1.7 USD per response (about 0.6 to 0.8 USD per single read). The low value assumes one human read per response. Replace with your own contract rates. |
What it leaves out: Gross avoided rater cost only. It leaves out engine licensing, validation studies per question, monitoring, rescoring and appeals, and the cost of errors, which can be large in reputation and student outcomes even when rare.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Texas Education Agency
United States · Government and public sector · 2023
The Texas Education Agency scores the short and extended constructed responses on the English language STAAR tests with a hybrid model: an automated scoring engine gives every response its first score, and at least 25 percent of responses per grade and subject are routed to trained human raters to monitor the engine. Responses with condition codes (for example blank, off topic, another language or vocabulary unlike the training data) or low confidence also go to humans, and a human score is always the score of record. The engine must agree with human raters as often as humans agree with each other before use; STAAR Spanish responses are scored only by humans. TEA says it used the hybrid approach to score all constructed responses in the December 2023 administration, after a study that rescored spring 2023 responses with the engine. TEA calls it an automated scoring engine and says it differs from AI that teaches itself: it is programmed on about 3,000 human scored field test responses per item.
No outcome disclosed.
ETS
United States · Education · 2012
ETS evaluates GRE Analytical Writing essays on a six point holistic scale, which includes a score from its own automated scoring engine, whose features ETS says are the result of nearly two decades of natural language processing research. All essays are also reviewed by trained analysts with essay similarity detection software and by experienced content experts. ETS cites a 2010 study (Attali, Bridgeman and Trapani) in which the engine agreed with a human rater on the GRE Issue and TOEFL Independent tasks more closely than two independent human raters agreed with each other. An ETS research report records that the engine was first implemented as a check score for the revised GRE General Test in August 2012.
No outcome disclosed.
Massachusetts Department of Elementary and Secondary Education
United States · Government and public sector · 2024
Massachusetts scores MCAS essays with AI trained on human scored examples of each score point, with humans giving 10 percent of AI scored essays a second read. The problem surfaced in summer 2025, when preliminary results went to districts. In Lowell, the example NBC10 Boston reports, a teacher found that some of her third grade students' scores did not add up and the issue went to district leaders; district leaders notified DESE. The state's testing contractor, Cognia, found that roughly 1,400 essays (of about 750,000 MCAS essays statewide) had not received the correct scores, which DESE attributed to a temporary technical issue in the process; the essays were rescored, 145 districts were notified and district data was corrected in August. DESE points to the discrepancy period in which districts can report issues with preliminary results as a check on accuracy.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Rubrics and anchor responses approved by educators for every question
- Human scored responses for each question from field tests, enough to train and validate
- Student group information to test for differences between engine and human scores
Systems to integrate
- Test delivery platform that captures responses
- Human scoring platform for routed responses and second reads
- Results reporting to schools with a discrepancy and rescore process
Complexity: High
The engine is the smaller part. The work is psychometric: per question training and validation, agreement thresholds, fairness checks across student groups, confidence routing, daily monitoring and a rescoring process that schools and families trust.
- 1
Decide where machine scoring is appropriate
Use it for tasks scored for writing quality or a clear rubric, not for tasks where the correctness of claims or creative reasoning is what counts. Keep human scoring for any language or test version the engine has not been validated on; Texas, for example, scores its Spanish STAAR responses entirely by hand.
- 2
Validate per question
Require the engine to agree with human raters at the same rate human raters agree with one another, with a similar score distribution, on responses the engine has not seen, for every question.
- 3
Route by confidence and condition
Send low confidence, borderline and unusual responses to humans, plus a fixed random share of all responses as an ongoing check. Make the human score the score of record.
- 4
Check fairness
Compare engine and human scores for student groups (for example by language background and disability) and investigate systematic differences before release.
- 5
Build in a discrepancy window
Release preliminary results to schools with time to report issues and a clear rescoring route, and publish how the process works.
Guardrails
- Human rescoring of a fixed share of responses and of all low confidence or flagged responses
- Human score as the score of record whenever a human scores a response
- Per question validation against human agreement before live use
- Daily monitoring of engine and human agreement during the scoring window
- A discrepancy period and rescoring process for schools and families
KPIs to instrument
- Exact and adjacent agreement between engine and human scores per question
- Share of responses routed to humans, by reason
- Differences between engine and human scores by student group
- Rescore requests and changed scores after release
- Cost and time per scored response
Human in the loop
Educators set the rubric and anchor responses; trained human raters score every flagged, low confidence and sampled response; scoring directors monitor agreement daily; and schools can challenge preliminary scores before results are final.
Common failure modes
- Systematic scoring errors on a set of responses
- A technical issue makes the engine score a set of essays incorrectly, as with about 1,400 Massachusetts essays in 2025, which DESE attributed to a temporary technical issue in the process. Sample human reads across the score range and give schools a discrepancy window.
- Responses unlike the training data
- Writing styles, structures or vocabulary that are rare in the training responses can be scored less reliably. Route responses the engine flags as unusual to humans and check agreement for different student groups.
- Gaming the engine
- Answers written to exploit surface features an engine may reward, such as length or rubric vocabulary, can score higher than their content deserves. Flag unusual responses, keep a random share of human reads and review what drives scores.
- Scoring what the engine can see
- Rubrics drift toward surface features such as length and punctuation. Keep educators in charge of the rubric and review what drives scores.
- Loss of trust
- Families and teachers distrust machine scores when the process is opaque. Publish how scoring works and how to request a rescore.
What are the risks and rules?
EU AI Act
High risk
Annex III point 3(b): AI systems intended to be used to evaluate learning outcomes in educational and vocational training institutions at all levels are high risk. Scoring that determines access to an institution or the level of education a student will receive is also covered by points 3(a) and 3(c). Schools and exam bodies that use such a system have the deployer obligations of Article 26.
Guidance
- Annex III, high risk AI systems (point 3, education and vocational training) (European Union, Europe). Lists AI systems intended to evaluate learning outcomes as high risk, which brings the requirements of Chapter III, Section 2 (Articles 8 to 15) on risk management, data governance, logging and human oversight, and the deployer obligations of Article 26.
- Scoring Process for STAAR Constructed Responses (Texas Education Agency, North America). A published description of a hybrid scoring process, including engine qualification against human agreement, confidence and condition code routing and the human score as the score of record.
- Generative artificial intelligence (AI) in education (UK Department for Education, Europe). The department's position on generative AI tools in schools and colleges, including its funding for AI tools that aim to reduce the burden of feedback and marking on teachers; relevant when teachers use AI to mark classroom work.
Controls to put in place
- Published description of the scoring process and how to request a rescore
- Per question validation records and agreement thresholds
- Fairness analysis across student groups before each release
- Logging of engine scores, confidence values and routing decisions
- Contract terms requiring the scoring vendor to report and correct errors
When it went wrong elsewhere
- AI grading issue affects hundreds of MCAS essays in Massachusetts. In 2025 the state's testing contractor Cognia found that roughly 1,400 MCAS essays (of about 750,000 statewide) had not received the correct scores under AI scoring, which DESE attributed to a temporary technical issue in the process. District leaders, including in Lowell, raised the problem after preliminary results were released, and the essays were rescored.
Frequently asked questions
- Is AI already used to score state tests?
- Yes. Massachusetts uses AI to score MCAS essays, trained on human scored examples of each score point, with humans giving 10 percent of AI scored essays a second read. Texas scores STAAR written responses with an automated scoring engine first and routes at least 25 percent, plus low confidence and unusual ones, to human raters; TEA says its engine is not AI in the sense of a system that teaches itself, but is programmed on about 3,000 human scored responses per question.
- How accurate is AI essay scoring?
- For suitable tasks, engines can agree with human raters as closely as two humans agree with each other. Texas requires this before an engine is used, and ETS cites a 2010 study in which its engine's agreement with a human rater on the TOEFL Independent and GRE Issue tasks was higher than between two human raters. Errors still happen: in 2025 about 1,400 Massachusetts essays did not receive the correct scores and were rescored.
- Is AI essay scoring high risk under the EU AI Act?
- Yes. Annex III point 3(b) lists AI systems intended to evaluate learning outcomes as high risk, which brings the requirements of Chapter III, Section 2 (Articles 8 to 15) on risk management, data governance, logging and human oversight. Schools and exam bodies that use such a system also have the deployer obligations of Article 26.
How to cite this page
Blits.ai AI Use Case Library, "AI scoring of essays and written answers in assessments", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/automated-scoring-of-written-responses. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published