AI use case

AI for public consultation response analysis

AI that reads every free text response to a public consultation or rulemaking comment period, proposes themes, maps each response to the themes that officials have validated, flags duplicates, campaign letters and responses that need special attention, and produces counts and summaries for the analysts who write the government's response.

By Len Debets · Last verified 27 September 2026 · 5 public deployments

At least 92%
Reported accuracy
Department for Transport, organization claim.
About 200,000
Interactions handled
Department for Transport (organization claim).
EUR 150,000 to EUR 1.5 million
Indicative value per year
A national ministry that runs 20 public consultations a year with substantial free text. Worked example, see how it is calculated.

What problem does it solve?

Governments ask the public for views before they change policy or make rules, and the answers arrive as free text: a few hundred responses to a technical consultation, or tens of thousands when an issue catches public attention. Every response has to be read, coded against a set of themes and counted, so that officials can show what people said and how it shaped the decision. Done by hand this can take months and, according to the UK Department for Transport, typically consumes over half of the consultation budget; the UK government notes that the work is often outsourced to contractors.

Speed is not the only problem. Coding is subjective, so two analysts can put the same response under different themes, and for very large consultations teams sometimes analyse a sample instead of every response. Mass campaigns and duplicate letters distort counts, and fake submissions have been used to manufacture the appearance of public support. Whatever tool is used, the government has to be able to show that every voice was heard and that the analysis was fair.

How does it work?

  1. Load and clean the responses. Responses from the consultation platform, email and regulations.gov style dockets are loaded per question, with personal data masked and exact and near duplicates (campaign letters) grouped so they are counted but read once.
  2. Propose themes. A language model, or an ensemble of models, reads the responses to each open question and proposes a set of themes, including rare but important points, with example responses for each.
  3. Validate the theme set with people. Analysts read a random sample of responses, merge, split, rename and add themes. Only the validated theme set is used from here on.
  4. Map every response. The model assigns each response to one or more validated themes and records its stance (agree, disagree, neutral) where the question asks for one. Responses that are off topic, abusive or that raise safeguarding concerns are flagged for a human.
  5. Check and report. Analysts review a sample of the mapping, correct errors and use the counts, summaries and representative quotes (always verified against the original response) to write the consultation response.
Audience
Back office
Autonomy
Copilot
Adoption
Early adopters
Channels
Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI for public consultation response analysis
KPIMedianReported rangeData pointsClaimed by
Interactions handledNot pooled
2000 to 200,000
22 organization
AccuracyToo few to pool
at least 92%
11 organization
Cost savingsNot pooled
about GBP 500,000
11 organization
Hours savedNot pooled
about 15,000 hours
11 organization

Value drivers: Employee productivity, Speed and cycle time, Lower cost to serve, Compliance quality.

Indicative value

A national ministry that runs 20 public consultations a year with substantial free text

EUR 150,000 to EUR 1.5 million

Consultation analysis cost avoided per year

How this is calculated

Formula: consultations * costPerConsultation * savingShare. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Consultations with free text analysis per year consultations, consultations per year1030Editorial assumption. The UK Department for Transport runs around 55 a year; replace with your own portfolio.
Cost of analysing and reporting one medium sized consultation costPerConsultation, EUR per consultation50,000100,000Editorial assumption, informed by the Department for Transport's estimate of GBP 80,000 to 100,000 for a medium sized consultation (about 15,000 responses). Replace with your own staff or contractor cost. Source
Share of that cost saved savingShare, fraction of cost0.30.5Conservative against the Department for Transport's modelled estimate of 50 to 70 percent for a notional medium sized consultation, because theme review, synthesis and report writing remain human work and small consultations save less. Source

What it leaves out: Gross analysis cost avoided only. It leaves out the cost of running and assuring the tool, the value of faster policy decisions, and the option of analysing every response in consultations that are sampled today.

Who already uses it?

5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Department for Science, Innovation and Technology (Incubator for Artificial Intelligence)

United Kingdom · Government and public sector · 2025

PilotGrade B

Consult is a generative AI tool built by the UK government's Incubator for Artificial Intelligence as part of the Humphrey suite. It proposes themes for each open question of a consultation, maps every response to those themes and shows the result in a dashboard that officials review and correct. Its first live use was on a Scottish Government consultation about non surgical cosmetic procedures, where it analysed more than 2,000 responses while officials also reviewed every response by hand; the government reported an F1 score of 0.76 in this first live evaluation. In July 2026 access was managed through a waitlist ahead of a cross government rollout planned for 2027.

  • Interactions handled: at least 2000, first live consultation (Scottish Government, six open questions)
    "Reviewing comments from over 2,000 consultation responses using generative AI, Consult identified key themes that feedback fell into across each of six qualitative questions."
    Claimed by: organization

Department for Transport

United Kingdom · Government and public sector · 2025

PilotGrade B

The Department for Transport runs around 55 consultations a year and co developed a Consultation Analysis Tool (CAT) with The Alan Turing Institute. An ensemble of large language models proposes themes and rare "golden insights", humans validate the themes in a structured review of a random sample, and the models then map every response to the validated themes. The published evaluation compares the tool with human coded datasets in blind and live settings, tests for accuracy differences across demographic groups (no evidence of systematic bias on the three questions analysed; small differences by ethnicity, under 3 percentage points and favouring minority groups, were of low practical significance) and estimates that the tool would save around 50 to 70 percent of the cost and time of a notional medium sized consultation (a modelled estimate, not a measured result).

  • Interactions handled: about 200,000, live pilots to date (December 2025 report)
    "The CAT has now been piloted on multiple live consultations, analysing 200,000 responses (exceeding 8 million words) to date."
    Claimed by: organization
  • Hours saved: about 15,000 hours, cumulative over four consultation and call for evidence or ideas projects to date (December 2025 report), modelled counterfactual
    "even with this investment, the CAT has roughly saved 15,000 hours of work to date compared to a scenario where all responses for all consultations are rigorously analysed by humans manually."
    Claimed by: organization
  • Cost savings: about GBP 500,000, cumulative over four consultation and call for evidence or ideas projects to date (December 2025 report), modelled counterfactual
    "the CAT has analysed 200,000 responses and over 8 million words, saving an estimated £0.5 million compared to a scenario where all responses for all consultations were rigorously analysed by humans."
    Claimed by: organization
  • Accuracy: at least 92%, theme mapping, raw agreement with human coders in blind and live evaluations
    "The CAT-vs-human inter-rater reliability (IRR), using metrics commonly employed in qualitative research to assess how consistently two or more researchers analyse the same data, achieved over 92% overall raw agreement in both our blind and non-blind evaluation designs."
    Claimed by: organization

U.S. Department of Transportation, Office of the Secretary

United States · Government and public sector · 2025

ProductionGrade B

The Office of the Secretary of Transportation reports a Public Comment Analyzer, deployed in February 2025, that uses a language model to categorise public comments by topic, detect their sentiment, generate summaries and provide daily updates on the comments to subject matter experts. The stated aim is to reduce the human effort of reading every comment and to let experts go from summaries to the underlying comments where needed. No measured outcome is published.

No outcome disclosed.

Centers for Disease Control and Prevention

United States · Government and public sector · 2023

ProductionGrade B

CDC reports in the 2025 federal AI use case inventory that it uses generative AI to analyse public comments on its proposed rules. For each comment the system records a stance (support, oppose or neutral), topics and sentiment, which regulatory analysts use when they review and summarise the feedback for the rulemaking record. The inventory lists the use case as deployed since July 2023. No outcome figures are published.

No outcome disclosed.

Board of Governors of the Federal Reserve System

United States · Government and public sector · 2021

ProductionGrade B

The Federal Reserve Board processes public comments on rulemakings, information collections and other proposals in its Comment Review System. The system uses traditional natural language processing for summaries, matching comments to lists of topics, entity identification and similarity matching, and flags duplicate and near duplicate comment letters. The Board states that all public comments are still reviewed in their entirety and that summaries only assist the review. The inventory lists it as deployed since July 2021; no outcome figures are published.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Exported responses per question, with respondent type (individual, organisation) where collected
  • A few previously human coded consultations to evaluate against
  • A policy on how personal data in responses is masked and retained

Systems to integrate

  • Consultation platform or regulations.gov style docket export
  • Email inbox for responses submitted outside the platform
  • Analysis workspace or dashboard for analysts

Complexity: Low

The data is text the government already holds and no transaction systems are touched. The work is in method, not integration: a defensible human review step, an evaluation against human coded samples and a bias check across respondent groups.

  1. 1

    Evaluate on consultations you have already coded

    Run the tool on two or three past consultations that humans coded, blind, and compare theme recall and mapping agreement with the human result before you use it live. Publish the method, as the UK Department for Transport did.

  2. 2

    Design the human theme review

    Decide how many responses analysts read per question before they sign off the theme set, and write it down. The Department for Transport reports 1 to 5 hours per 100 responses for this step.

  3. 3

    Run the first live consultation in parallel

    On the first live use, let analysts also code every response by hand, as the Scottish Government did with Consult, and measure where the two disagree.

  4. 4

    Handle campaigns and duplicates explicitly

    Group identical and near identical responses, count them, and report organised campaigns separately so that one template letter does not read as thousands of independent views.

  5. 5

    Check for bias across respondent groups

    Compare mapping accuracy for responses from different groups (for example by writing style, language or respondent type) and act on any gap before the method is used at scale.

Guardrails

  • Only human validated themes are used for mapping and counts
  • Every quote in the report is copied from the original response, never from a model summary
  • Personal data is masked before responses reach a model and in stored outputs
  • The tool never decides policy or weights responses; it organises them for analysts
  • Duplicate and campaign detection is reported, not used to discard responses

KPIs to instrument

  • Theme recall and mapping agreement against a human coded sample, per question
  • Analyst hours per 1,000 responses, before and after
  • Days from consultation close to published response
  • Share of mapped responses changed by analysts during review
  • Accuracy differences across respondent groups

Human in the loop

Analysts own the theme framework, review a random sample of the mapping for every question and write the response. Policy officials see the underlying responses behind every theme. Responses flagged for safeguarding or abuse go to a named person.

Common failure modes

Missing the rare but important point
Models favour frequent themes and can miss a single expert response that changes the policy. Ask for rare themes explicitly and have analysts read a random sample.
Hallucinated or softened quotes
A summary that paraphrases respondents can put words in their mouths. Pull quotes from the source text only.
Counting as if consultations were polls
Consultation respondents are self selected; reporting that a percentage of respondents felt something invites misreading. Report counts with context, as the Department for Transport cautions.
Campaigns and fake submissions distorting results
Mass template letters or fabricated submissions can swamp genuine views. Detect duplicates and unusual submission patterns and report them openly.

What are the risks and rules?

EU AI Act

Minimal risk

Organising and summarising consultation responses for analysts does not decide on individuals and is not listed in Annex III, so no high risk obligations apply. If AI generated text is published to inform the public on matters of public interest without human review and editorial responsibility, Article 50(4) requires disclosure.

Guidance

Controls to put in place

  • Published method and evaluation before live use, including a transparency record
  • Documented human theme review with a sample size rule
  • Audit trail from each reported count back to the underlying responses
  • Bias testing across respondent groups on every major model or prompt change
  • Retention and masking rules for personal data in responses

When it went wrong elsewhere

Frequently asked questions

How accurate is AI at coding consultation responses?
Close to human coders when humans validate the themes. The UK Department for Transport reports over 92% raw agreement between its tool and human coders (an F1 score of 0.75 for theme mapping in its blind evaluation), and the UK government reported an F1 score of 0.76 on Consult's first live consultation. Human review of the theme set remains necessary; without it, the Department for Transport's tool found about 75% of the human themes.
Can AI replace the analysts?
No. It replaces most of the reading and tagging, while analysts decide the themes, check the mapping and write the response. The Department for Transport estimates, in a model rather than a measured result, savings of 50 to 70% of the cost of a medium sized consultation, with review, synthesis and report writing remaining human work.
Should respondents be told that AI analyses their responses?
Yes. Say so in the consultation document and privacy notice, publish the method and keep a human accountable for the analysis. In the UK the Algorithmic Transparency Recording Standard is mandatory for government departments, and the Department for Science, Innovation and Technology has published a record for Consult.

How to cite this page

Blits.ai AI Use Case Library, "AI for public consultation response analysis", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/public-consultation-response-analysis. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryRetail and ecommerce

AI for voice of the customer and feedback analysis

AI that reads every piece of free text customer feedback, such as survey verbatims, NPS comments, reviews, social posts, chat and call transcripts, and turns it into themes, sentiment, drivers and suggested actions that a named owner can act on, so the organization hears all of its customers instead of a sample.

Deployments
5 public, best grade B
Reported accuracy
84%
SBF Group, vendor claim
Government and public sector

AI drafting copilot for civil servants for correspondence, briefings and ministerial replies

A generative AI assistant that drafts replies to correspondence from the public and elected representatives, briefings, submissions and summaries for civil servants, grounded in the department's approved lines, policy documents and case data, with the official editing and approving every word before it is sent or cleared.

Deployments
5 public, best grade B
Autonomy
Copilot
Cross industryBanking

AI for complaints root cause and systemic issue analysis

AI that reads the free text of complaints across all channels, clusters them into themes, separates systemic causes from one off events, links each theme to the product, process or control behind it and routes the insight to the owner who can fix it, with a human validating every root cause and every remediation.

Deployments
3 public, best grade B
Autonomy
Copilot
Government and public sector

AI for freedom of information request processing

AI that helps a public body handle freedom of information and open government requests: logging and clarifying requests, spotting duplicates, searching and deduplicating the records in scope, proposing redactions with the exemption that applies, and drafting the response letter, with an FOI officer deciding what is released.

Deployments
4 public, best grade B
Autonomy
Copilot