AI use case

AI assistant for tax questions and filing support

An AI assistant that answers taxpayers' questions about taxes, deadlines, refunds and payments, lets authenticated taxpayers check their status or set up a payment plan within set rules, and guides them through filing, while assessments, penalties and disputes stay with the tax authority's staff and systems.

By Len Debets · Last verified 26 September 2026 · 4 public deployments

83%
Reported accuracy
HM Revenue and Customs, organization claim.
20%
Reported contact deflection
HM Revenue and Customs, organization claim.
USD 1.2 million to USD 7.2 million
Indicative value per year
A national tax authority with 3 million assisted contacts a year. Worked example, see how it is calculated.

What problem does it solve?

Tax authorities face sharply seasonal demand. Around filing deadlines and after every policy change, phone lines and webchat fill with the same questions: where is my refund, how do I get a reference number, can I pay in instalments, what does this notice mean. Long waits push people to give up, file late or file wrongly, which creates more work later in compliance and correspondence.

The questions are repetitive but the stakes are not trivial: a wrong answer about a deadline or a relief can cost the taxpayer money. That is why the tax authority assistants on this page (HMRC's digital assistant and the IRS voice bots and chatbots) classify intent and return approved content, and add authenticated actions such as payment plans only behind identity checks.

How does it work?

  1. Recognise the intent. The assistant classifies the question (refund status, payment plan, notice, registration, how to file) on chat or on the phone.
  2. Answer from approved content. General questions get answers written or approved by the tax authority, with links to the guidance; unclear questions get a choice of likely meanings.
  3. Authenticate for account questions. For refund status, balance or a payment plan, the taxpayer verifies identity (shared secrets, PIN or national login) before the assistant reads the account.
  4. Act within rules. Within fixed limits the assistant can set up or change a payment plan or grant a payment extension, and confirms the result.
  5. Escalate. Complex, disputed or personal situations go to a human adviser, who sees the whole conversation and completes identity checks.
Audience
Customer facing
Autonomy
Supervised agent
Adoption
Mainstream
Channels
Web chat, Phone and voice, Mobile app, WhatsApp

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI assistant for tax questions and filing support
KPIMedianReported rangeData pointsClaimed by
Interactions handledNot pooled
3 million to 5.5 million
22 organization
AccuracyToo few to pool
83%
11 organization
Contact deflectionToo few to pool
20%
11 organization
Users servedNot pooled
at least 200,000
11 vendor

Value drivers: Lower cost to serve, Customer experience, Compliance quality, Inclusion and access.

Indicative value

A national tax authority with 3 million assisted contacts a year

USD 1.2 million to USD 7.2 million

Human handled contact cost avoided per year

How this is calculated

Formula: contacts * routineShare * containment * costPerContact. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Assisted phone and chat contacts per year contacts, contacts per year3,000,0003,000,000The reference authority.
Share of contacts on routine topics (refund status, payments, forms) routineShare, fraction of contacts0.40.6Editorial assumption. Replace with your own contact reason data.
Share of routine contacts the assistant resolves containment, fraction of routine contacts0.20.4Editorial assumption. For context, HMRC reports that webchat escalations to advisers fell 20% while assistant interactions grew 18.8%.
Cost of a human handled contact costPerContact, USD per contact510Editorial assumption. Replace with your own fully loaded cost.

What it leaves out: Gross contact cost only. It leaves out the effect on filing accuracy and late payment, the value of shorter queues at peak and the cost of building and running the assistant.

Who already uses it?

4 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

HM Revenue and Customs

United Kingdom · Government and public sector · 2025

ScaledGrade B

HMRC's digital assistant answers tax questions typed in plain language, matching them to intents with natural language understanding and replying with non personalised answers that link to GOV.UK guidance. It covers 60 of HMRC's more than 120 taxes, asks the user to choose between likely meanings when unsure, and needs no login. Complex questions escalate to webchat with a human adviser, who sees the whole prior conversation and only continues after identity and verification checks.

  • Interactions handled: at least 5.5 million, tax year 2024/25 to 6 March 2025
    "The digital assistant has had 5.48m interactions in the tax year 2024/25 (to date as at 6 March 2025)."
    Claimed by: organization
  • Accuracy: 83%, NLU test set, March 2025, known intents
    "83.03% on known intents"
    Claimed by: organization
  • Contact deflection: 20%, webchat escalations to an adviser, tax year 2024/25 to 6 March 2025
    "For webchat, 909,000 users have escalated to speak to an adviser in the tax year 2024/25 (to date as at 6 March 2025). That is a 20% decrease on the 2023/24 tax year."
    Claimed by: organization

Internal Revenue Service

United States · Government and public sector · 2022

ScaledGrade B

Since 2021 the IRS has put intent based voice bots on many toll free lines: payment plans and balance due (with authentication so taxpayers can set up or change a payment plan), Where's My Refund and amended return status, notice clarifications, Economic Impact Payments and the Advance Child Tax Credit. On IRS.gov, chatbots answer FAQs on refunds, identity theft, payments and relief. The inventory stresses that answers are not generated: the model classifies the question and returns content approved by the business owner, or routes the call to a live assistor.

  • Interactions handled: at least 3 million, calls answered by voice bots, cumulative to June 2022
    "To date, the voice bots have answered over 3 million calls."
    Claimed by: organization

Internal Revenue Service

United States · Government and public sector · 2020

ProductionGrade B

The IRS Linguistic Policy, Tools and Services team uses a cloud machine translation application on AWS, with the IRS Publication 850 glossary of English and Spanish tax terms, to translate text and files between English and Spanish, Chinese, Korean and Vietnamese and speed up responses to taxpayers. Separately, staff use SYSTRAN neural translation, augmented with a domain dictionary and translation memories, to triage non English documents for relevance to case work and as a starting point for manual translation. Both appear in the federal AI inventory as in operation.

No outcome disclosed.

ClearTax

India · Technology and software · 2024

ProductionGrade C

ClearTax, an Indian online tax filing platform, built a generative AI agent on WhatsApp, using Azure OpenAI models, for low income and blue collar workers who have tax deducted at source from their income but rarely file a return, and so miss refunds and lack the income proof lenders ask for. The agent explains taxes and filing, walks users through the filing workflow and supports nine Indian vernacular languages. It is a private filing service, not a tax authority channel.

  • Users served: at least 200,000, workers who filed their returns independently
    "200,000+ blue-collar workers successfully filed ITRs independently"
    Claimed by: vendor
  • Customer savings: about INR 300 million, tax refunds claimed by users, cumulative
    "300 million INR unlocked in tax refunds, improving access to credit opportunities"
    Claimed by: vendor

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Approved answers and guidance per tax and topic, with owners and effective dates
  • Contact reason data by season, to pick intents and plan for peaks
  • Rules for payment plans and extensions that can be applied without judgment

Systems to integrate

  • Taxpayer account, refund and payment systems through APIs
  • Identity verification and step up authentication
  • Telephony IVR and webchat platforms with handover to advisers
  • Notice and correspondence systems, to explain letters by reference

Complexity: Medium

Information answers are low complexity; authenticated actions such as payment plans need integration with taxpayer account systems, strong identity checks and strict rules on what the assistant may change.

  1. 1

    Start where the queues are

    Use contact reason data to pick the few intents that dominate peak season (refund status, payments, notices), as the IRS did when it put voice bots on its Economic Impact Payment, notice and payment lines.

  2. 2

    Keep answers approved

    Let the model classify and retrieve, and serve content the business owner approved. The IRS inventory stresses that its bots do not generate answers.

  3. 3

    Add authenticated self service

    Add account questions and payment plans behind identity checks, with limits (amount, term) written as rules, and confirm every change back to the taxpayer.

  4. 4

    Design escalation with context

    Pass the conversation to the adviser, as HMRC's webchat advisers see the assistant history, and complete identity checks before personal discussion.

  5. 5

    Prepare for peaks and changes

    Retest before each filing season and after each budget, and plan capacity for the spike.

Guardrails

  • No assessment, penalty or dispute outcome is decided by the assistant
  • Account data only after identity verification, at the level the action needs
  • Actions such as payment plans only within written limits, with confirmation to the taxpayer
  • Answers from approved content with effective dates; refusal when the topic is out of scope
  • Personal data and tax identifiers masked in logs and model prompts

KPIs to instrument

  • Containment per intent, counting repeat contacts within seven days as not contained
  • Escalations to advisers, before and after launch
  • Intent recognition accuracy on a labelled test set
  • Payment plans set up through the assistant and their default rate
  • Satisfaction on assistant and adviser conversations

Human in the loop

Advisers handle escalations, disputes, hardship and anything outside the rules. Content owners approve every answer and every change after a budget; a team samples conversations weekly and reviews failed intents and complaints.

Common failure modes

Wrong deadline or relief answers
A fluent but wrong answer costs the taxpayer money and trust. Serve approved content, show effective dates and refuse outside scope.
Peak season collapse
The assistant is tested in quiet months and fails at the deadline. Load test and retest before each season.
Authentication friction
Taxpayers fail identity checks and fall back to the phone. Offer several proportionate methods and measure drop off.
Private helpers without oversight
Third party filing assistants (such as ClearTax's WhatsApp agent) help people file but are not the authority; make official guidance easy for them to use and keep the authority's own channel authoritative.

What are the risks and rules?

EU AI Act

Limited risk (transparency)

A taxpayer assistant must tell people they are interacting with an AI system (Article 50). It is not listed in Annex III as long as it only informs and applies fixed rules. It becomes high risk under Annex III point 5(a) if it evaluates eligibility for, or grants, reduces, revokes or reclaims, public assistance benefits (which can include benefits paid through the tax system). Recital 59 says systems used for administrative proceedings by tax and customs authorities are not high risk law enforcement systems; audit selection and risk scoring are covered on a separate page.

Guidance

Controls to put in place

  • AI disclosure and a route to a human adviser on every channel
  • Entry in the public AI inventory or transparency register
  • Written limits for every action the assistant can take on an account
  • Change control tied to the budget and filing season calendar
  • Retention limits and access control on conversation logs containing tax data

Frequently asked questions

How much volume do tax assistants handle?
HMRC's digital assistant had 5.48 million interactions in the 2024/25 tax year up to 6 March 2025, and IRS voice bots had answered over 3 million calls by June 2022, about a year after the first one went live in May 2021.
Do tax authorities use generative AI for these answers?
Not in the deployments on this page. The IRS inventory states its chatbots and voice bots return content predetermined by content owners, and HMRC matches intents to approved answers. The IRS tested a generative AI chatbot for volunteer tax preparers as a proof of concept (listed as retired in 2025), while the private filing service ClearTax runs a WhatsApp agent on Azure OpenAI models.
How accurate is intent recognition?
HMRC reports 83.03% accuracy on known intents in its March 2025 test set. Plan for the rest with disambiguation questions and an easy route to a human.

How to cite this page

Blits.ai AI Use Case Library, "AI assistant for tax questions and filing support", last verified 26 September 2026, https://www.blits.ai/ai-use-cases/tax-questions-and-filing-assistant. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Government and public sector

AI assistant for citizen information and government services

An AI assistant that answers residents' and businesses' questions about government services in plain language, grounded only in official guidance with links to the source, points them to the right online service or office, and hands anything personal, urgent or outside its content to a human with the context attached.

Deployments
9 public, best grade B
Reported accuracy
at least 76%
Foreign, Commonwealth and Development Office, organization claim
Government and public sector

AI for tax compliance risk scoring and audit selection

Models that score tax returns, taxpayers and transactions for the risk of error, underreporting or fraud, so that a tax administration spends its audit and compliance capacity where the risk is highest, with an officer deciding every compliance action and the selection itself monitored for fairness.

Deployments
3 public, best grade B
Autonomy
Assist
Government and public sector

AI assistant for benefits eligibility questions and applications

An AI assistant that helps people understand which public benefits and grants may apply to them, explains the rules and documents in plain language, guides them through the application and checks it for completeness, while the eligibility decision stays with the agency's rules and caseworkers.

Deployments
7 public, best grade B
Reported accuracy
97%
Department for Work and Pensions, organization claim
Government and public sector

AI translation and interpretation for multilingual public services

AI that translates government content, documents and conversations between officials and the public, in writing and in real time speech, so people can use public services in their own language, with human translators and interpreters reviewing what carries legal or safety weight.

Deployments
8 public, best grade B
Autonomy
Copilot
Cross industryBanking

AI agent for first line contact centre service

An AI agent that answers the first line of inbound customer contact on phone, chat and messaging, resolves general and routine questions end to end in the customer's own language, and routes everything complex, sensitive or regulated to the right human team with the context attached.

Deployments
18 public, best grade B
Median containment rate
47%
7 deployments