What problem does it solve?
When a critical system degrades, the monitoring estate fires a burst of alerts across applications, infrastructure and dependent services. Engineers spend the first part of the incident working out which alerts belong together, which team owns the problem and what changed, while business and customer teams ask for updates. Mizuho and IBM described the pattern plainly: when an error is detected, operators receive an influx of messages and reports, which makes it hard to pinpoint the cause and delays recovery.
Time to recover is not only a cost question. For regulated firms an outage of a critical service can become a reportable event. DORA requires EU financial entities to detect, manage, record and classify ICT related incidents, to report major ones and to review them afterwards. NIS2 sets incident reporting duties for essential and important entities, including telecom operators. APRA CPS 230 expects Australian regulated entities to keep critical operations within tolerance levels through severe disruptions and to notify APRA of tolerance breaches. The knowledge that shortens an incident (runbooks, past post incident reviews, recent changes) exists, but it is scattered and nobody has time to search it at 3 a.m.
How does it work?
- Correlate. Alerts, logs, traces and change events are grouped into one probable incident using topology and timing, so responders see one problem instead of a stream of separate alerts.
- Route. A classifier or a set of team agents decides which team owns the incident, based on service ownership and past incidents, and pages them with the correlated evidence.
- Suggest causes. The AI ranks recent changes and known failure patterns as likely root causes and retrieves the matching runbook steps and similar past incidents, with links so the engineer can verify each suggestion.
- Propose, do not execute. Remediation is proposed as a concrete command or change; a policy layer checks it, and an engineer confirms it before anything runs.
- Communicate. The AI drafts status updates for stakeholders at a set cadence from the incident timeline, for the incident commander to approve.
- Learn. After recovery it drafts the post incident review from the timeline, chat and changes, and proposes follow up actions and runbook updates.
- Audience
- Employee facing
- Autonomy
- Copilot
- Adoption
- Early adopters
- Channels
- Internal tools, Microsoft Teams
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Accuracy | 90% | 42% to 98% | 3 | 2 organization, 1 vendor |
| Time to repair reduction | Too few to pool | 20% to 38% | 2 | 1 organization, 1 vendor |
| Detection improvement | Too few to pool | 25% | 1 | 1 vendor |
Value drivers: Speed and cycle time, Risk and loss reduction, Employee productivity, Customer experience.
Indicative value
A bank with 120 major IT incidents a year on customer facing services
USD 26,880 to USD 187,200
Engineering time released during and after major incidents per year
How this is calculated
Formula: incidents * (hoursToRestore * reduction * responders + reviewHoursSaved) * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Major incidents per year incidents, incidents per year | 120 | 120 | The reference organization. Replace with your own count of priority 1 and 2 incidents. |
| Average hours from detection to mitigation hoursToRestore, hours per incident | 2 | 4 | Editorial assumption. Replace with your own mean time to restore. |
| Reduction in time to mitigate reduction, fraction of time | 0.1 | 0.25 | Conservative against the evidence on this page (Microsoft reports a 38% time to mitigate reduction for one team; TD Bank's vendor reports 20% faster response), because those are the best early results. |
| Engineers engaged per incident responders, engineers | 4 | 8 | Editorial assumption including the incident commander and service owners. |
| Fully loaded engineer hour hourlyCost, USD per hour | 80 | 120 | Editorial assumption. Replace with your own rate. |
| Hours saved drafting each post incident review and status updates reviewHoursSaved, hours per incident | 2 | 5 | Editorial assumption; the review still needs the owning team's analysis. |
What it leaves out: Engineering time only. It leaves out the largest effect, the revenue, customer harm and regulatory exposure avoided by restoring service faster, which depends on the service and is best estimated per critical business service against its impact tolerance.
Who already uses it?
5 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
United States · Technology and software · 2026
Google site reliability engineers use an agent in the Gemini CLI across an outage: reading the page, investigating, proposing mitigations, finding the root cause and drafting the postmortem. Every proposed change passes a policy layer (for example rules that need two person approval) and a forced human confirmation, and every proposal and approval is logged. No outcome figures are published.
No outcome disclosed.
Microsoft
United States · Technology and software · 2025
Azure uses the Triangle System to triage incidents with AI agents. In local triage, one agent per engineering team, built on the team's historical incidents and troubleshooting guides, accepts or rejects an incoming incident on the team's behalf and can recommend the team it should move to; a global triage layer coordinates the agents to route incidents. Local triage has been in production since mid 2024 and was live for six teams in January 2025, with more than 15 onboarding; Microsoft reports triage accuracy and a time to mitigate reduction for one team as initial results.
- Accuracy: 90%, initial results, as of January 2025
"The initial results are promising, with agents achieving 90% accuracy and one team saw a reduction in their TTM of 38%, significantly reducing the impact to customers."
Claimed by: organization - Time to repair reduction: 38%, one team, initial results
"The initial results are promising, with agents achieving 90% accuracy and one team saw a reduction in their TTM of 38%, significantly reducing the impact to customers."
Claimed by: organization
Meta
United States · Technology and software · 2024
Meta's reliability investigation tooling uses a heuristic retriever (code ownership, the runtime code graph of impacted systems) to narrow thousands of recent code changes to a few hundred, then a fine tuned Llama 2 model ranks them to the five most likely root causes when an investigation is opened. Meta reports the result from backtesting on historical investigations in its web monorepo and stresses that responders must be able to verify the suggestions.
- Accuracy: 42%, backtesting on historical investigations, web monorepo
"Based on exhaustive backtesting, with historical investigations and the information available at their start, 42% of these investigations had the root cause in the top five suggested code changes."
Claimed by: organization
Mizuho Financial Group
Japan · Banking · 2024
Mizuho and IBM ran a three month proof of concept that added patterns likely to cause errors in incident response to generative AI on IBM watsonx and linked it to the application that supports event detection, so that operators flooded with messages during a disruption can find the cause faster. Accuracy was measured on actual data. Both firms said they planned to expand the proof of concept and apply it to production, and to use generative AI for incident management and failure analysis next.
- Accuracy: 98%, three month trial
"The new solution demonstrated a 98% accuracy[1] in monitoring and responding to error messages during a three-month trial."
Claimed by: vendor
TD Bank
Canada · Banking · 2024
TD Bank is consolidating roughly ten monitoring tools across its clouds into one observability platform whose AI identifies the root cause of emerging issues across its hybrid, multicloud estate, so teams can respond to transaction failures before users are affected. The customer story does not describe the AI as generative, and the outcome figures come from the vendor.
- Detection improvement: 25%
"As a result, TD Bank is identifying 25% more incidents proactively and responding to them 20% faster."
Claimed by: vendor - Time to repair reduction: 20%
"As a result, TD Bank is identifying 25% more incidents proactively and responding to them 20% faster."
Claimed by: vendor
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- A current service catalog with owners and dependencies
- Alerts, logs and traces from the observability platform
- Change records and deployment history linked to services
- Past incidents and post incident reviews with the confirmed root cause
- Runbooks with an owner and a review date
Systems to integrate
- Observability and monitoring platforms
- IT service management tool for incidents, problems and changes
- Paging and on call scheduling
- ChatOps in Microsoft Teams or Slack
- Code repositories and deployment pipelines for change history
Complexity: Medium
Summaries and drafted updates are quick wins. Correlation and root cause suggestions depend on a service map, clean change records and labeled past incidents; without them, suggestions are generic. Automated remediation is a separate, higher risk step that most firms defer.
- 1
Start with the write ups, not the fixes
The lowest risk value is drafting stakeholder updates and post incident reviews from the incident timeline. It earns trust with engineers and builds the labeled incident history the later steps need.
- 2
Clean the service map and change feed
Correlation and routing are only as good as ownership data and change records. Fix the services with the most incidents first.
- 3
Add root cause suggestions with evidence
Show a short ranked list, each item with the change, the log excerpt or the past incident that supports it. Meta narrows its suggestions to the top five code changes, measures the ranker by backtesting on historical investigations and holds back low confidence answers.
- 4
Measure on history before going live
Replay past incidents and measure how often the true cause was in the suggestions and how often routing picked the right team. Publish the number to the engineers who will use it.
- 5
Keep remediation behind confirmation
Let the AI propose commands, pass them through a policy layer (no global restarts at peak, two person approval for high impact changes) and require an engineer to confirm. Google's SRE tooling logs what the AI proposed and what the human approved.
- 6
Close the loop
Feed confirmed root causes and runbook fixes from each review back into the knowledge base, so the next incident starts with better context.
Guardrails
- No change to production without explicit confirmation by an authorized engineer
- A policy layer that blocks or escalates high impact commands regardless of what the AI proposes
- Every suggestion shows its evidence (change, log, past incident) so it can be verified in seconds
- Low confidence suggestions are held back rather than shown
- Full logging of AI proposals, human decisions and timestamps into the incident timeline
- Secrets and customer data masked before logs reach a model
KPIs to instrument
- Mean time to detect, to engage the right team and to mitigate, before and after, per service
- Share of incidents where the confirmed root cause was among the AI suggestions
- Routing accuracy (incidents that stayed with the first team paged)
- Alerts per incident after correlation
- Time from resolution to a published post incident review
Human in the loop
The incident commander owns the incident, decides on customer communication and approves every status update. Engineers authorize every remediation. The owning team signs off the post incident review and its actions; the AI drafts, it does not conclude.
Common failure modes
- Plausible but wrong root cause
- An engineer anchors on a confident suggestion and loses time. Show evidence per suggestion, hold back low confidence ones and track the hit rate openly.
- Automation that amplifies the outage
- An automated fix runs against the wrong target or at the wrong time. Keep confirmation and a policy layer in front of every mutation, and rehearse in game days.
- Stale runbooks retrieved with authority
- The AI surfaces an outdated procedure. Give every runbook an owner and a review date and prefer recent post incident reviews.
- Noise moved, not removed
- Correlation merges unrelated alerts or hides a second incident. Let responders split incidents easily and review correlation quality weekly.
What are the risks and rules?
EU AI Act
Minimal risk
An internal tool that supports engineers on IT incidents; it is not a use listed in Annex III and makes no decisions about people. Annex III point 2 covers AI used as a safety component in the management and operation of critical digital infrastructure, and recital 55 limits safety components to systems that directly protect the physical integrity of that infrastructure or the health and safety of persons and property. A triage copilot that proposes causes and fixes to engineers does not normally do that, but operators of critical digital infrastructure (cloud, data centers, telecom networks) should confirm this for their own design.
Guidance
- Digital Operational Resilience Act (Regulation (EU) 2022/2554) (European Union, Europe). Financial entities must detect, manage, record and classify ICT related incidents (Articles 17 and 18), report major ones (Article 19) and review major incidents afterwards (Article 13); AI drafted timelines and reviews become part of that record.
- Operational risk management (CPS 230) (Australian Prudential Regulation Authority, Asia Pacific). Regulated entities must keep critical operations within tolerance levels through severe disruptions and notify APRA of operational risk incidents likely to have a material impact and of disruptions to a critical operation outside tolerance, which is the measure AIOps value should be judged against.
Controls to put in place
- Written limits on what the AI may propose and what always needs a named approver
- Incident timeline that records AI suggestions, human decisions and timestamps
- Backtest results per release of the model or prompts before engineers rely on it
- Periodic review of suggestion hit rate and routing accuracy by the SRE or operations lead
- Inventory entry for the AIOps system with an owner and a review date
Frequently asked questions
- How accurate is AI root cause analysis today?
- It helps, but it is not an oracle. Meta reports that in backtesting 42% of investigations had the root cause in the top five suggested code changes, and Microsoft reports 90% accuracy for its team triage agents in early results. Treat suggestions as a ranked shortlist with evidence, not an answer.
- Should the AI fix incidents on its own?
- Not at first, and not for high impact changes. Keep the AI proposing and an engineer confirming, with a policy layer in between; Google's SRE tooling forces a confirmation step and logs what the AI proposed and what the human approved. Automate only narrow, reversible fixes after a track record.
- Where does the value show up first?
- In routing and write ups. Getting the incident to the right team faster and drafting updates and post incident reviews saves time on every incident, while root cause suggestions improve as the labeled incident history grows.
- Does this matter to regulators?
- Yes for financial firms. DORA in the EU sets rules for managing, recording and reporting ICT related incidents, and CPS 230 in Australia expects critical operations to stay within tolerance levels, so the AI's suggestions and the human decisions belong in the incident record.
How to cite this page
Blits.ai AI Use Case Library, "AI for IT incident triage and root cause analysis (AIOps)", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/aiops-incident-triage. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published