What problem does it solve?
Periodic control testing checks a sample of items at a point in time, with evidence gathered from control owners by email and screenshots. The first line collects the evidence, the second line reviews it, and the result describes the control as it was when the sample was drawn. A control that fails soon after a test can go unnoticed until the next cycle, and is found only if the failing items happen to be in the sample.
Risk and control self assessments have the same weakness. Anaptyss, a services vendor, describes them as heavily manual, with control narratives that differ across teams and business units. Rules raise the bar at the same time: APRA CPS 230 requires regulated entities to regularly monitor, review and test controls for design and operating effectiveness and to report the results to senior management.
- Anaptyss, a services vendor, states that risk and control self assessments in banks often take three to six weeks to complete, with some complex assessments extending beyond eight weeks.From 20 Days to 5: The Operational Economics of AI-Led RCSA Execution (2026)
How does it work?
- Codify the control. Each automatable control gets a testable rule (for example: every payment above a threshold has a second approver who is not the initiator) and a list of the evidence that proves it.
- Collect evidence continuously. Agents pull records, logs, approvals and documents from source systems on a schedule, instead of asking control owners for screenshots.
- Read the artefacts. Document AI reads policies, sign off records, reconciliations and reports and maps each artefact to the control and the period it covers.
- Test the whole population. Every item is tested against the rule; exceptions are grouped and explained with the evidence attached.
- Human judgement on exceptions. A tester or control owner reviews each exception, decides whether it is a control failure and records the reason.
- Prepare the RCSA. The AI drafts control narratives and proposed ratings from test results, incidents and losses; the business reviews, challenges and owns the final assessment.
- Audience
- Back office
- Autonomy
- Supervised agent
- Adoption
- Emerging
- Channels
- Internal tools
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Interactions handled | Not pooled | at least 29,000 | 1 | 1 organization |
Value drivers: Compliance quality, Risk and loss reduction, Employee productivity, Speed and cycle time.
Indicative value
A bank that tests 2,000 key controls a year
USD 96,000 to USD 1.1 million
Control testing effort released per year
How this is calculated
Formula: controls * hoursPerTest * automatable * effortSaved * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Key controls tested per year controls, controls per year | 2,000 | 2,000 | The reference bank. Replace with your own control library. |
| Hours per control test (evidence collection, testing, review) hoursPerTest, hours per control test | 10 | 20 | Editorial assumption across first and second line effort. Replace with your own time records. |
| Share of controls suitable for automated testing automatable, fraction of controls | 0.2 | 0.4 | Editorial assumption; system based controls automate first, judgement based controls stay manual. |
| Share of test effort saved on those controls effortSaved, fraction of test effort | 0.4 | 0.7 | Editorial assumption, replace with your own pilot results. No public bank benchmark was found. |
| Fully loaded cost of a control tester hourlyCost, USD per hour | 60 | 100 | Editorial assumption, replace with your own. |
What it leaves out: Effort only. It leaves out the build and integration cost, the value of finding failures months earlier across the full population, and the saving in audit and regulatory findings.
Who already uses it?
3 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Federal Deposit Insurance Corporation
United States · Government and public sector · 2025
The FDIC's Division of Finance reports a use case in pre deployment that uses classical machine learning to monitor financials and invoices for proper submittal, duplicate payments and other abnormalities, described as an automated assist for its auditing and monitoring work. For each gap it would show why it was flagged and possible causes, in visual tables with drill downs to contract numbers and agency sections. Not yet live; no results published.
No outcome disclosed.
Pension Benefit Guaranty Corporation
United States · Government and public sector · 2025
The Pension Benefit Guaranty Corporation reports a use case in pre deployment that applies AI to IT security and privacy control assessments: evaluating controls against federal cybersecurity and privacy guidelines, working through the supporting evidence, generating findings and drafting control implementation statements. The stated aims are a higher volume of control assessments, less manual work and shorter control tailoring and implementation times. It is not yet live and no results are published.
No outcome disclosed.
U.S. Department of the Interior
United States · Government and public sector · 2024
The Interior Department's Office of Grants Management built an in house set of AI tools, operational since April 2024, for three reviews of financial assistance awards that had been manual, inconsistent across bureaus and labour intensive: project descriptions, pre award eligibility validations in SAM.gov and budget submissions. The department says it needed a way to conduct internal controls testing, eligibility checks and budget reviews at scale. The tools produce automated scoring, flags for risks or inconsistencies, cross walks between budget documents and audit ready records aligned with internal control requirements. The 2024 inventory files the project description and SAM.gov work under internal controls testing, including a large language model that scores SAM.gov documents against the award date and a planned Azure app, to be built with Microsoft, that would let bureau staff test project descriptions themselves.
- Interactions handled: at least 29,000, per year, financial assistance actions
"Together, these outputs streamline oversight, strengthen regulatory compliance, and create a consistent, defensible documentation trail for more than 29,000 annual financial assistance actions."
Claimed by: organization
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- A control library with owners, objectives and testable attributes
- Access to transaction, approval, access and configuration data in source systems
- Incident, loss and issue data linked to controls
- Prior test results to compare automated and manual outcomes
Systems to integrate
- Governance, risk and compliance (GRC) platform with the control library and issues
- Core banking, payments, identity and access management and ticketing systems
- Document stores for policies, sign offs and reconciliations
- Incident and operational loss databases
Complexity: High
Testing logic is simple once a control is codified. The work is in rewriting controls so they are testable, reaching evidence in dozens of source systems and agreeing with audit and the regulator that automated tests are reliable.
- 1
Pick controls that are data rich and rule based
Start with access, approval, reconciliation and segregation of duties controls whose evidence already sits in systems. Leave judgement heavy controls for later.
- 2
Rewrite each control as a test
For every control in scope, write the rule, the population, the evidence and the exception definition, and have the control owner and second line sign it off.
- 3
Run parallel with manual testing
For one or two cycles test both ways, compare results and explain every difference before retiring the manual test. Share the comparison with internal audit.
- 4
Industrialise exception handling
Route exceptions to owners with evidence, reasons and due dates, and link confirmed failures to issues and the RCSA.
- 5
Draft, never auto approve, the RCSA
Use AI to prepare narratives and proposed ratings from data, and require the business to review and change them, so the assessment stays theirs.
Guardrails
- Every exception and every proposed rating is reviewed and decided by a human
- Test rules are versioned and signed off by the control owner and second line
- The agent reads source systems with read only access through an allow list
- Every test run stores its population, evidence and result for audit
- Automated tests are recalibrated when the underlying process or system changes
KPIs to instrument
- Share of key controls tested on the full population
- Time from control failure to detection
- Exceptions raised, confirmed as failures and closed on time
- Hours of evidence collection and testing per control, before and after
- Differences between automated and manual test outcomes during parallel runs
Human in the loop
Control owners and testers adjudicate every exception, second line approves test designs, the business owns the self assessment ratings and internal audit reviews the reliability of the automated testing.
Common failure modes
- Testing the wrong population
- The agent tests the data it can reach, not the full population the control covers. Reconcile populations to source totals every run.
- Rubber stamped RCSA
- The business accepts AI drafted ratings without challenge and the self assessment loses its point. Require recorded review and track how often drafts change.
- Silent test decay
- A system change breaks the extraction and the test passes on empty data. Alert on volume anomalies and zero populations.
- Exception floods
- A poorly specified rule produces thousands of exceptions that nobody reviews. Tune on parallel runs and group exceptions by cause.
What are the risks and rules?
EU AI Act
Depends on design
Testing controls over transactions and systems is not an Annex III use. Controls that monitor and evaluate individual employees' behaviour, such as trading or access conduct, can fall under Annex III point 4(b), so the design decides the tier.
Guidance
- Revisions to the principles for the sound management of operational risk (Basel Committee on Banking Supervision, Global). The Basel Committee's principles for operational risk management and the control environment, revised in 2021 with updated guidance on change management and ICT.
- Operational risk management (CPS 230) (Australian Prudential Regulation Authority, Asia Pacific). Australia's cross industry operational risk standard for APRA regulated entities, covering operational risk controls, critical operations and material service providers.
- Article 4, AI literacy (European Union, Europe). Providers and deployers must take measures to support the AI literacy of staff who use and oversee AI systems (the amended wording asks them to support it rather than ensure a sufficient level). Here that means training testers on the limits of automated results.
Controls to put in place
- Inventory entry for the testing AI with an owner, scope and validation status
- Signed off test specifications per control, under change control
- Population reconciliation and completeness checks on every run
- Parallel run evidence retained before a manual test is retired
- Internal audit review of the reliability of automated testing
Frequently asked questions
- Does continuous controls testing replace control testers?
- It changes their work. Evidence collection and routine testing move to automation, and testers spend their time on exceptions, control design and judgement based controls that cannot be codified.
- Which controls should be automated first?
- Controls whose evidence already sits in systems and whose rule can be written precisely: access reviews, approvals and limits, reconciliations and segregation of duties. Judgement heavy controls stay manual or AI assisted rather than automated.
- Is there public evidence this works?
- Public bank figures are scarce. The US Department of the Interior reports AI tools in production since 2024 that run internal controls testing and produce audit ready records for more than 29,000 financial assistance actions a year. The FDIC and the Pension Benefit Guaranty Corporation list similar tools as pre deployment in the 2025 federal AI inventory. Vendor claims about faster RCSA cycles exist but are not tied to named banks.
How to cite this page
Blits.ai AI Use Case Library, "AI for continuous controls testing and control self assessment", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/continuous-controls-testing. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published