What problem does it solve?
A security operations centre lives on alerts: from the SIEM, the endpoint and email protection tools, identity systems and cloud platforms, plus the suspicious emails that staff report with a button in their mail client. Many of these alerts are false positives (St. Luke's University Health Network describes finding the true threats "amidst a sea of false positives"), but each one has to be opened, enriched and judged, because the one real intrusion hides among them.
The judgment is not the slow part; the legwork is. An analyst checks the sender and the links, looks up the IP address and the file hash, pulls the sign in history of the user, compares the event with earlier incidents and writes up what they found. Before AI, that legwork took up to 20 minutes per event at SEP2 and up to 30 minutes per alert at Human Managed, and 10 to 77 minutes per identity investigation at Avanade. Most alerts take far less, but at St. Luke's triaging hundreds of alerts a day still took hours every day. At that volume the queue wins: alerts wait, analysts burn out and real threats are found late.
Security teams also struggle with consistency. No two analysts document an investigation the same way, which makes handovers between shifts, escalations to forensics and reports to leadership slower than they need to be.
How does it work?
- Pick up the alert. The agent is triggered by a new alert in the SIEM or extended detection and response (XDR) console, or by an email a user reported as suspicious.
- Gather the evidence. It queries the connected tools the way an analyst would: email headers, URLs and attachments, endpoint and firewall logs, identity sign ins, threat intelligence on indicators, and similar past incidents.
- Reason to a verdict. It classifies the alert (malicious, suspicious, benign, false positive) with a confidence level and writes down the evidence and the reasoning behind it, so an analyst can check the verdict in a minute instead of rebuilding it.
- Act within limits. Clear false positives, such as a reported marketing email, are closed with the reasoning attached. Anything malicious or uncertain is escalated with a draft incident summary; containment actions such as isolating a device or disabling an account are proposed, and an analyst approves them.
- Learn from feedback. Analysts correct wrong verdicts in plain language, and the agent's instructions, allow list and examples are updated, so the same mistake is not repeated.
- Audience
- Employee facing
- Autonomy
- Supervised agent
- Adoption
- Early adopters
- Channels
- Internal tools, Microsoft Teams, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Productivity gain | 60% | 60% to 70% | 3 | 1 organization, 2 vendor |
| Handling time reduction | Too few to pool | 97% | 1 | 1 vendor |
| Hours saved | Not pooled | about 200 hours | 1 | 1 organization |
| Productivity gain | Too few to pool | 20x | 1 | 1 vendor |
| Time to repair reduction | Too few to pool | Not pooled: up to 80% | 0plus 1 up to | 1 organization |
Value drivers: Employee productivity, Risk and loss reduction, Speed and cycle time, Compliance quality.
Indicative value
An enterprise security operations centre that triages 50,000 to 100,000 alerts and reported emails a year
USD 45,000 to USD 1 million
Analyst triage time released, valued at cost per year
How this is calculated
Formula: alertsPerYear * minutesPerAlert / 60 * timeSavedShare * costPerHour. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Alerts and reported emails triaged by analysts per year alertsPerYear, alerts per year | 50,000 | 100,000 | Editorial assumption for a large enterprise SOC. Replace with the count from your SIEM or case management system. |
| Average analyst minutes per alert before AI minutesPerAlert, minutes per alert | 3 | 10 | An average across the alert mix, well below the upper bounds on this page (up to 20 minutes at SEP2 and up to 30 minutes at Human Managed for a single alert), because most alerts are closed quickly. St. Luke's says triaging hundreds of alerts a day took hours, and its nearly 200 hours saved a month across thousands of closed false positives implies a few minutes per reported email. Even the high case (about 10,000 hours a year) is roughly four times St. Luke's 2,400 hours a year from phishing alone, for a SOC that covers every alert type. |
| Share of triage time the agent saves timeSavedShare, fraction of triage time | 0.3 | 0.6 | Conservative against the evidence on this page, because a first deployment covers only part of the alert types. Google Cloud reports that manual triage at Human Managed was reduced by more than 60%. TÜV SÜD reports analysis about 60% to 70% faster; if faster means more analyses per hour, that is roughly 40% less time per analysis, and if it means 60% to 70% less time, it sits above this range. |
| Fully loaded cost of a SOC analyst hour costPerHour, USD per hour | 60 | 100 | Editorial assumption, replace with your own fully loaded cost or your managed security provider's rate. |
What it leaves out: Counts analyst time released on triage only. It leaves out the cost of the AI and the security tools it calls, the integration work, the value of finding real incidents sooner and the cost of a wrong verdict, which the guardrails on this page are meant to keep rare.
Who already uses it?
7 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
U.S. Immigration and Customs Enforcement
United States · Government and public sector · 2024
ICE reports in the 2025 federal AI use case inventory that its security operations centre uses an AI Assisted Compromise Email Detector (AACED) to review a collection of emails exchanged with Microsoft under Emergency Directive 24-02. Named entity recognition flags personal data and keywords, and a chat interface lets analysts ask questions with an email as context, so they can find indicators of compromise faster. This was a review of a fixed set of emails tied to one directive rather than ongoing alert triage, and the inventory classifies it as classical machine learning rather than an agent, so it is a partial fit for this use case. The inventory lists it as deployed since June 2024 and gives no measured result.
No outcome disclosed.
Federal Housing Finance Agency
United States · Government and public sector · 2021
FHFA reports in the 2025 federal AI use case inventory that it uses KnowBe4 PhishER to manage the suspicious emails its staff report. Machine learning classifies each email as spam, phishing or malicious and sends the response automatically, so a cybersecurity specialist does not have to analyse hundreds of reported emails by hand. It is a machine learning classifier rather than a generative AI agent, so it covers the classification step of this use case only. The inventory lists the use case as deployed since April 2021 with risk reduction as its benefit, and says the tool reduces the time to respond to staff, but gives no figures.
No outcome disclosed.
Avanade
United States · Professional services · 2026
Avanade embedded Microsoft Security Copilot in its identity and security investigations, run with Microsoft Entra. After a five week pilot in which engineers turned recurring investigation patterns into prompt based workflows, it applied the same approach to security operations and built custom agents that consolidate signals, recommendations and playbook steps for analysts. Most of the reported gains are in identity incident handling, closer to identity and access support than to SOC alert triage: guest account investigations dropped from 10 minutes to 1 minute. A separate study of more than 4,100 security operations investigations reports gains in speed, quality and error rates.
- Time to repair reduction: up to 80%, identity incident resolution time
"By turning expert knowledge into repeatable, AI-assisted workflows, we reduced resolution times by up to 80% and enabled the team to take on more work with greater consistency and confidence."
Claimed by: organization - Productivity gain: 70%, speed and efficiency across more than 4,100 security operations investigations
"In a separate study of more than 4,100 security operations investigations, Avanade found a 70% improvement in speed and efficiency, a 7% improvement in investigation and documentation quality, and a 7% reduction in human error."
Claimed by: vendor
SEP2
United Kingdom · Technology and software · 2026
SEP2, a UK managed security provider, built triage and threat intelligence agents on Google's Gemini Enterprise Agent Platform alongside Google Security Operations. When an alert fires, the agents gather data from firewalls, endpoints and threat feeds in about a minute, a task that took analysts up to 20 minutes, and hand a standardised summary to the human team, who review it and authorise remediation. Engineers also use Gemini to write detection rules from plain language.
- Productivity gain: 20x, gathering the data for a new security alert before human review, from up to 20 minutes to one minute
"20x faster to triage new security alerts"
Claimed by: vendor
Human Managed
Singapore · Technology and software · 2025
Human Managed, a Singapore based security intelligence provider for large enterprises, uses Google Security Operations, Vertex AI and Gemini to analyse its customers' alerts and logs. Before, its specialists needed up to 30 minutes per alert to analyse the data and send a contextual alert to the customer; now statistical models and generative AI produce a confidence score with an explanation, and customers get prioritised alerts with remediation advice within 15 minutes.
- Handling time reduction: 97%, time to triage an alert, from up to 30 minutes to under one minute
"With the system monitoring data and continuously learning about new threats, manual triage has been reduced by more than 60% and the time to triage alerts has been slashed to less than one minute – a 97% reduction that helps Human Managed's customers respond to security incidents much faster."
Claimed by: vendor - Productivity gain: at least 60%, reduction in manual triage work
"With the system monitoring data and continuously learning about new threats, manual triage has been reduced by more than 60% and the time to triage alerts has been slashed to less than one minute – a 97% reduction that helps Human Managed's customers respond to security incidents much faster."
Claimed by: vendor
St. Luke's University Health Network
United States · Healthcare · 2025
St. Luke's University Health Network, a US hospital network with more than 23,000 employees, runs Microsoft Security Copilot across Defender and Sentinel. Its Security Alert Triage Agent (formerly the Phishing Triage Agent) reads user reported emails, decides whether each is a genuine phishing attempt or a false alarm, explains its verdict in plain text and closes false positives on its own, so analysts can move to threat hunting. Copilot also drafts incident reports that analysts edit and escalate.
- Hours saved: about 200 hours, per month, phishing alert triage
"It’s saving us nearly 200 hours monthly by autonomously handling and closing thousands of false positive alerts."
Claimed by: organization
TÜV SÜD
Germany · Professional services · 2024
TÜV SÜD, the German testing and certification group, protects about 28,000 employees with Microsoft Defender solutions, runs Microsoft Sentinel as its SIEM and joined the early adopter programme for Security Copilot. Its analysts use Copilot inside Defender to enrich alerts, investigate threats, start remediation and produce consistent investigation reports, and the company says new analysts become effective within months.
- Productivity gain: about 60%, speed of threat analysis, reported as 60% to 70% faster
"We analyze results about 60% to 70% faster with Security Copilot."
Claimed by: organization
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Alerts and reported emails available through the SIEM, XDR or email security API
- Read access to logs from endpoints, identity, email, firewalls and cloud platforms
- Threat intelligence feeds for indicators such as IP addresses, domains and file hashes
- Past incidents with their outcomes, to calibrate verdicts and write examples
- Written triage runbooks per alert type
Systems to integrate
- SIEM and XDR platform (alerts, queries, incident updates)
- Email security gateway and the user reporting mailbox
- Identity provider for sign in history and account actions
- Endpoint detection and response for device context and isolation
- Security orchestration (SOAR) or case management for tickets and playbooks
- Collaboration tool such as Microsoft Teams for analyst notifications
Complexity: Medium
The model work is modest; the integrations and the trust are the work. The agent needs read access to many security tools, a clear policy on what it may close on its own, and weeks of side by side running before analysts stop double checking every verdict.
- 1
Start with the noisiest, best understood alert type
User reported phishing is the usual first choice: high volume, mostly benign, and the evidence (headers, links, attachments, sender history) is well defined. Measure the current volume, handling time and false positive share before you begin.
- 2
Write the triage runbook down
Turn the senior analysts' practice into explicit steps and decision rules per alert type, including which evidence decides the verdict. The agent follows this, and it is what you audit against later.
- 3
Connect read only first
Give the agent read access to the tools it needs and nothing else. Let it produce verdicts and summaries next to the analysts for several weeks without closing anything.
- 4
Compare verdicts and calibrate
Compare the agent's verdicts with the analysts' on the same alerts, per alert type. Only where agreement is consistently high, and the misses are understood, allow the agent to close false positives on its own.
- 5
Keep humans on containment
Isolating devices, disabling accounts, blocking senders and deleting emails stay proposed actions that an analyst approves, with the agent's evidence attached, until the organization has a documented reason to automate a specific action.
- 6
Extend alert by alert
Add alert types one at a time (identity, endpoint, data loss prevention), each with its own runbook, test set and agreement threshold, and keep sampling closed alerts every week. Alert types that watch individual employees, such as data loss prevention and insider risk, need a fresh legal and works council review first.
Guardrails
- The agent closes only the alert types and verdicts it has been cleared for; everything else goes to an analyst
- Containment and account actions require analyst approval, logged with the evidence behind them
- Content from emails, attachments and web pages is treated as untrusted input, so instructions hidden in a phishing email cannot steer the agent
- Read only, least privilege service accounts for every connected tool
- Every verdict carries its evidence and reasoning, stored with the incident record
KPIs to instrument
- Mean time to triage per alert type, before and after
- Share of alerts closed by the agent, and the share of those later reopened
- Agreement rate between agent and analyst verdicts on a weekly sample
- Missed true positives found in sampling or later investigations
- Analyst hours spent on triage versus threat hunting and response
Human in the loop
Analysts own every containment decision and every escalation. They review a random sample of alerts the agent closed each week, correct wrong verdicts, and approve each new alert type or automated action before it goes live. The agent prepares; the analyst decides.
Common failure modes
- A confident false negative
- The agent closes a real phishing email or intrusion as benign. Limit autonomous closure to well calibrated alert types, sample closed alerts every week and treat every miss as an incident with a root cause.
- Prompt injection through the evidence
- Attackers write text into emails or files that tells the model to mark them safe. Separate instructions from evidence, and test the agent with adversarial samples before and after every change.
- Automation without the runbook
- The agent is switched on without written triage rules, so nobody can say whether a verdict was right. Write the runbook first; it is also what auditors and new analysts need.
- Analysts stop looking
- Once the agent is usually right, reviews become rubber stamps. Keep a structured sample review and rotate who does it.
What are the risks and rules?
EU AI Act
Depends on design
Triage of phishing, endpoint, network and cloud alerts for an organization's own cyber defence is not listed in Annex III. Recital 55 of the AI Act says that components intended to be used solely for cybersecurity purposes should not qualify as safety components, so the agent does not fall under Annex III point 2 (critical infrastructure), and for this scope the tier is minimal. The design changes that when the agent triages identity, data loss prevention, insider risk or user behaviour alerts in a way that scores or monitors individual employees: monitoring and evaluating the behaviour of persons in a work relationship falls under Annex III point 4(b), so that scope needs its own high risk assessment before it goes live. The Article 50(1) duty to disclose AI interaction does not apply because it is obvious to a reasonably well informed analyst that they are working with an AI agent. An operator that lets AI act autonomously on network or operational technology controls should assess that design separately, and reading employees' emails and sign in data remains subject to data protection law.
Rules that apply
Guidance
- NIST SP 800-61 Rev. 3, Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile (NIST, North America). The US reference for incident detection, analysis and response, mapped to the NIST Cybersecurity Framework 2.0; a useful baseline for the runbooks an agent follows.
- OWASP Top 10 for LLM Applications (OWASP, Global). Lists prompt injection and excessive agency among the main risks of LLM applications, both directly relevant when an agent reads attacker controlled emails and can act in security tools.
Controls to put in place
- Inventory entry for the agent with an accountable owner, the alert types in scope and the actions it may take
- Written triage runbook per alert type, versioned with the agent's instructions
- Audit trail of every verdict, the evidence used and every action proposed or taken
- Weekly sample review of closed alerts with tracked agreement and misses
- Adversarial test set (prompt injection, look alike domains, novel lures) run on every change
Frequently asked questions
- How much analyst time does AI alert triage save?
- It depends on the alert mix and on how much the agent may close on its own. St. Luke's University Health Network says its triage agent saves nearly 200 hours a month on reported phishing, Google Cloud reports that manual triage at Human Managed was reduced by more than 60%, and TÜV SÜD reports analysis about 60% to 70% faster. These are vendor case studies, so measure your own baseline before you plan on similar numbers.
- Should the AI close alerts by itself?
- Only for alert types where its verdicts have matched your analysts' for weeks, and usually only for false positives such as benign reported emails. St. Luke's describes the agent closing thousands of false positive alerts, with analysts focusing on the real threats it surfaces. Containment actions such as isolating a device should stay with an analyst.
- Can attackers trick a triage agent?
- Yes, that is the specific risk here: the agent reads text the attacker wrote. Treat email and file content as untrusted data, keep the agent's permissions read only, and test it with prompt injection and novel phishing samples on every change.
- Is phishing triage a separate use case from SOC alert triage?
- No, it is a common first step in the same job: user reported phishing is high in volume and its evidence is well defined. Microsoft has renamed its Phishing Triage Agent the Security Alert Triage Agent, and the US Federal Housing Finance Agency has automated responses to user reported suspicious emails since 2021.
How to cite this page
Blits.ai AI Use Case Library, "AI for security alert triage and investigation in the SOC", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/security-alert-triage-and-investigation. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published