AI use case

AI for continuous controls testing and control self assessment

AI that moves control testing from periodic samples to continuous, full population assurance: it collects evidence from source systems, maps each artefact to the control it supports, tests every transaction or record against the control's rule, flags exceptions for a human to judge and prepares the risk and control self assessment from incident and loss data for the business to review.

By Len Debets · Last verified 27 September 2026 · 3 public deployments

At least 29,000
Interactions handled
U.S. Department of the Interior (organization claim).
USD 96,000 to USD 1.1 million
Indicative value per year
A bank that tests 2,000 key controls a year. Worked example, see how it is calculated.

What problem does it solve?

Periodic control testing checks a sample of items at a point in time, with evidence gathered from control owners by email and screenshots. The first line collects the evidence, the second line reviews it, and the result describes the control as it was when the sample was drawn. A control that fails soon after a test can go unnoticed until the next cycle, and is found only if the failing items happen to be in the sample.

Risk and control self assessments have the same weakness. Anaptyss, a services vendor, describes them as heavily manual, with control narratives that differ across teams and business units. Rules raise the bar at the same time: APRA CPS 230 requires regulated entities to regularly monitor, review and test controls for design and operating effectiveness and to report the results to senior management.

How does it work?

  1. Codify the control. Each automatable control gets a testable rule (for example: every payment above a threshold has a second approver who is not the initiator) and a list of the evidence that proves it.
  2. Collect evidence continuously. Agents pull records, logs, approvals and documents from source systems on a schedule, instead of asking control owners for screenshots.
  3. Read the artefacts. Document AI reads policies, sign off records, reconciliations and reports and maps each artefact to the control and the period it covers.
  4. Test the whole population. Every item is tested against the rule; exceptions are grouped and explained with the evidence attached.
  5. Human judgement on exceptions. A tester or control owner reviews each exception, decides whether it is a control failure and records the reason.
  6. Prepare the RCSA. The AI drafts control narratives and proposed ratings from test results, incidents and losses; the business reviews, challenges and owns the final assessment.
Audience
Back office
Autonomy
Supervised agent
Adoption
Emerging
Channels
Internal tools

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

Value benchmarks for AI for continuous controls testing and control self assessment
KPIMedianReported rangeData pointsClaimed by
Interactions handledNot pooled
at least 29,000
11 organization

Value drivers: Compliance quality, Risk and loss reduction, Employee productivity, Speed and cycle time.

Indicative value

A bank that tests 2,000 key controls a year

USD 96,000 to USD 1.1 million

Control testing effort released per year

How this is calculated

Formula: controls * hoursPerTest * automatable * effortSaved * hourlyCost. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Key controls tested per year controls, controls per year2,0002,000The reference bank. Replace with your own control library.
Hours per control test (evidence collection, testing, review) hoursPerTest, hours per control test1020Editorial assumption across first and second line effort. Replace with your own time records.
Share of controls suitable for automated testing automatable, fraction of controls0.20.4Editorial assumption; system based controls automate first, judgement based controls stay manual.
Share of test effort saved on those controls effortSaved, fraction of test effort0.40.7Editorial assumption, replace with your own pilot results. No public bank benchmark was found.
Fully loaded cost of a control tester hourlyCost, USD per hour60100Editorial assumption, replace with your own.

What it leaves out: Effort only. It leaves out the build and integration cost, the value of finding failures months earlier across the full population, and the saving in audit and regulatory findings.

Who already uses it?

3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Federal Deposit Insurance Corporation

United States · Government and public sector · 2025

AnnouncedGrade B

The FDIC's Division of Finance reports a use case in pre deployment that uses classical machine learning to monitor financials and invoices for proper submittal, duplicate payments and other abnormalities, described as an automated assist for its auditing and monitoring work. For each gap it would show why it was flagged and possible causes, in visual tables with drill downs to contract numbers and agency sections. Not yet live; no results published.

No outcome disclosed.

Pension Benefit Guaranty Corporation

United States · Government and public sector · 2025

AnnouncedGrade B

The Pension Benefit Guaranty Corporation reports a use case in pre deployment that applies AI to IT security and privacy control assessments: evaluating controls against federal cybersecurity and privacy guidelines, working through the supporting evidence, generating findings and drafting control implementation statements. The stated aims are a higher volume of control assessments, less manual work and shorter control tailoring and implementation times. It is not yet live and no results are published.

No outcome disclosed.

U.S. Department of the Interior

United States · Government and public sector · 2024

ProductionGrade B

The Interior Department's Office of Grants Management built an in house set of AI tools, operational since April 2024, for three reviews of financial assistance awards that had been manual, inconsistent across bureaus and labour intensive: project descriptions, pre award eligibility validations in SAM.gov and budget submissions. The department says it needed a way to conduct internal controls testing, eligibility checks and budget reviews at scale. The tools produce automated scoring, flags for risks or inconsistencies, cross walks between budget documents and audit ready records aligned with internal control requirements. The 2024 inventory files the project description and SAM.gov work under internal controls testing, including a large language model that scores SAM.gov documents against the award date and a planned Azure app, to be built with Microsoft, that would let bureau staff test project descriptions themselves.

  • Interactions handled: at least 29,000, per year, financial assistance actions
    "Together, these outputs streamline oversight, strengthen regulatory compliance, and create a consistent, defensible documentation trail for more than 29,000 annual financial assistance actions."
    Claimed by: organization

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • A control library with owners, objectives and testable attributes
  • Access to transaction, approval, access and configuration data in source systems
  • Incident, loss and issue data linked to controls
  • Prior test results to compare automated and manual outcomes

Systems to integrate

  • Governance, risk and compliance (GRC) platform with the control library and issues
  • Core banking, payments, identity and access management and ticketing systems
  • Document stores for policies, sign offs and reconciliations
  • Incident and operational loss databases

Complexity: High

Testing logic is simple once a control is codified. The work is in rewriting controls so they are testable, reaching evidence in dozens of source systems and agreeing with audit and the regulator that automated tests are reliable.

  1. 1

    Pick controls that are data rich and rule based

    Start with access, approval, reconciliation and segregation of duties controls whose evidence already sits in systems. Leave judgement heavy controls for later.

  2. 2

    Rewrite each control as a test

    For every control in scope, write the rule, the population, the evidence and the exception definition, and have the control owner and second line sign it off.

  3. 3

    Run parallel with manual testing

    For one or two cycles test both ways, compare results and explain every difference before retiring the manual test. Share the comparison with internal audit.

  4. 4

    Industrialise exception handling

    Route exceptions to owners with evidence, reasons and due dates, and link confirmed failures to issues and the RCSA.

  5. 5

    Draft, never auto approve, the RCSA

    Use AI to prepare narratives and proposed ratings from data, and require the business to review and change them, so the assessment stays theirs.

Guardrails

  • Every exception and every proposed rating is reviewed and decided by a human
  • Test rules are versioned and signed off by the control owner and second line
  • The agent reads source systems with read only access through an allow list
  • Every test run stores its population, evidence and result for audit
  • Automated tests are recalibrated when the underlying process or system changes

KPIs to instrument

  • Share of key controls tested on the full population
  • Time from control failure to detection
  • Exceptions raised, confirmed as failures and closed on time
  • Hours of evidence collection and testing per control, before and after
  • Differences between automated and manual test outcomes during parallel runs

Human in the loop

Control owners and testers adjudicate every exception, second line approves test designs, the business owns the self assessment ratings and internal audit reviews the reliability of the automated testing.

Common failure modes

Testing the wrong population
The agent tests the data it can reach, not the full population the control covers. Reconcile populations to source totals every run.
Rubber stamped RCSA
The business accepts AI drafted ratings without challenge and the self assessment loses its point. Require recorded review and track how often drafts change.
Silent test decay
A system change breaks the extraction and the test passes on empty data. Alert on volume anomalies and zero populations.
Exception floods
A poorly specified rule produces thousands of exceptions that nobody reviews. Tune on parallel runs and group exceptions by cause.

What are the risks and rules?

EU AI Act

Depends on design

Testing controls over transactions and systems is not an Annex III use. Controls that monitor and evaluate individual employees' behaviour, such as trading or access conduct, can fall under Annex III point 4(b), so the design decides the tier.

Guidance

  • Revisions to the principles for the sound management of operational risk (Basel Committee on Banking Supervision, Global). The Basel Committee's principles for operational risk management and the control environment, revised in 2021 with updated guidance on change management and ICT.
  • Operational risk management (CPS 230) (Australian Prudential Regulation Authority, Asia Pacific). Australia's cross industry operational risk standard for APRA regulated entities, covering operational risk controls, critical operations and material service providers.
  • Article 4, AI literacy (European Union, Europe). Providers and deployers must take measures to support the AI literacy of staff who use and oversee AI systems (the amended wording asks them to support it rather than ensure a sufficient level). Here that means training testers on the limits of automated results.

Controls to put in place

  • Inventory entry for the testing AI with an owner, scope and validation status
  • Signed off test specifications per control, under change control
  • Population reconciliation and completeness checks on every run
  • Parallel run evidence retained before a manual test is retired
  • Internal audit review of the reliability of automated testing

Frequently asked questions

Does continuous controls testing replace control testers?
It changes their work. Evidence collection and routine testing move to automation, and testers spend their time on exceptions, control design and judgement based controls that cannot be codified.
Which controls should be automated first?
Controls whose evidence already sits in systems and whose rule can be written precisely: access reviews, approvals and limits, reconciliations and segregation of duties. Judgement heavy controls stay manual or AI assisted rather than automated.
Is there public evidence this works?
Public bank figures are scarce. The US Department of the Interior reports AI tools in production since 2024 that run internal controls testing and produce audit ready records for more than 29,000 financial assistance actions a year. The FDIC and the Pension Benefit Guaranty Corporation list similar tools as pre deployment in the 2025 federal AI inventory. Vendor claims about faster RCSA cycles exist but are not tied to named banks.

How to cite this page

Blits.ai AI Use Case Library, "AI for continuous controls testing and control self assessment", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/continuous-controls-testing. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryBanking

Generative AI copilot for internal audit

A copilot for internal auditors that drafts planning memos and document request lists from prior audits, summarises large evidence sets, builds risk and control matrices from policies and process documents, and drafts findings and reports, with every statement traceable to its evidence and a qualified auditor accountable for every conclusion.

Deployments
3 public, best grade C
Reported handling time reduction
55%
Banco Bradesco, vendor claim
Cross industryBanking

AI for policy drafting and policy gap analysis

An assistant that takes a new or changed obligation, finds every internal policy, standard and procedure it touches, flags clauses that now conflict or are silent, and drafts the updated wording in house style as a redline for the policy owner to approve.

Deployments
3 public, best grade B
Autonomy
Copilot
Cross industryBanking

AI for complaints root cause and systemic issue analysis

AI that reads the free text of complaints across all channels, clusters them into themes, separates systemic causes from one off events, links each theme to the product, process or control behind it and routes the insight to the owner who can fix it, with a human validating every root cause and every remediation.

Deployments
3 public, best grade B
Autonomy
Copilot
Cross industryBanking

AI for third party and vendor risk due diligence

AI that reviews a vendor's security questionnaires, SOC and assurance reports, contracts and model documentation against the organization's control requirements, researches the vendor's ownership, sanctions, financial health and adverse media, drafts the risk assessment for a human to approve and keeps the register of material service providers current with ongoing monitoring.

Deployments
4 public, best grade B
Autonomy
Copilot
Cross industryBanking

AI system and model inventory with shadow AI discovery

A governed register of every AI system and model an organization builds, buys or uses, with its owner, purpose, data, risk tier and approval status, kept current by AI that discovers unregistered use, reads the documentation and assembles the evidence a board, auditor or supervisor asks for.

Deployments
4 public, best grade B
Autonomy
Copilot
Cross industryBanking

AI regulatory horizon scanning and obligation mapping

An AI system that continuously reads publications from the regulators and standard setters an organization answers to, classifies each item by relevance and urgency, breaks new rules into individual obligations and maps them to the internal policies and controls that meet them, so compliance owners see what changed and where the gaps are.

Deployments
2 public, best grade B
Autonomy
Assist