AI use case

AI for software vulnerability triage and remediation

AI that takes security findings from scanners, fuzzers and bug reports, filters out duplicates and false positives, reproduces and ranks the real ones, and drafts a code fix with a test for each, which a developer reviews and merges through the normal change process.

By Len Debets · Last verified 27 September 2026 · 4 public deployments

USD 12,000 to USD 336,000
Indicative value per year
A software organization that fixes 2,000 to 5,000 security findings a year. Worked example, see how it is calculated.

What problem does it solve?

Security tools find far more problems than teams can fix. Static analysis (SAST), dependency scanners, fuzzers, penetration tests and bug bounty reports all feed the same backlog, and each finding needs someone to confirm it is real, work out how serious it is, find the owner and write a fix that does not break anything. Developers see security tickets as interruptions to feature work, so the backlog grows and old findings stay open.

AI is making the imbalance sharper on both sides. Models are now used to find vulnerabilities (Google's Big Sleep agent found bugs in Chrome's V8 engine in 2025). Separately, the Chrome security team reports receiving more bug reports by March 2026 than in all of 2025, without naming a single cause, and says triaging a single report used to take 5 to 30 or more minutes of expert time. At the same time, low quality AI generated reports waste maintainers' time, and developers using coding assistants ship more code, and so more findings, than before. Fixing has to scale as fast as finding, without letting unreviewed code into production.

How does it work?

  1. Collect and deduplicate. Findings arrive from scanners, fuzzers, CI pipelines and external reports. The system drops spam and duplicates and checks that each one describes a real security issue in scope.
  2. Validate and rank. Where possible it reproduces the bug (running the proof of concept or the failing test), checks whether the vulnerable code is reachable, and assigns a severity using written guidelines. Findings that mitigations already neutralise are marked for a human to confirm and close.
  3. Route to the owner. The finding goes to the right component and code owner with the stack trace, severity and context attached.
  4. Draft the fix. A fixing agent proposes one or more candidate patches; a separate critic step or reviewer model checks them against the style guide, and the fix is rebuilt and rescanned to prove the finding is gone. Test writing agents add a regression test.
  5. Human review and merge. The developer reviews the pull request like any other change and merges, edits or rejects it. Rejections and edits feed back into the prompts and examples.
Audience
Employee facing
Autonomy
Copilot
Adoption
Early adopters
Channels
Internal tools, API and system to system

What is it worth?

Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.

No public deployment has disclosed a measurable outcome yet.

Value drivers: Risk and loss reduction, Employee productivity, Speed and cycle time, Compliance quality.

Indicative value

A software organization that fixes 2,000 to 5,000 security findings a year

USD 12,000 to USD 336,000

Developer time released on security fixes, valued at cost per year

How this is calculated

Formula: findingsPerYear * hoursPerFix * aiFixShare * timeSavedShare * costPerHour. The low scenario uses every low input, the high scenario every high input.

InputLowHighBasis
Security findings fixed per year findingsPerYear, findings per year2,0005,000Editorial assumption for a mid sized engineering organization. Replace with the count from your vulnerability management or code scanning tool.
Developer hours per manual fix, including review hoursPerFix, hours per finding1.54GitHub reports a median of 1.5 hours to resolve code scanning alerts manually, measured in its public beta on alerts raised in pull requests on new code; the high end is an editorial allowance for older backlog findings, which take longer. Source
Share of findings where an AI drafted fix is accepted aiFixShare, fraction of findings0.10.2An editorial range around the only measured acceptance rate, Google's 15% of sanitizer bugs fixed by its LLM pipeline with human review. Chrome's candidate fixes for most vulnerabilities are drafts, not accepted fixes, so they are not used here. Source
Share of developer time saved on those findings timeSavedShare, fraction of fix time0.50.7GitHub reports a median of 28 minutes with Copilot Autofix against 1.5 hours manually, about 69% less, measured in its public beta on new alerts in pull requests; the low end allows for backlog fixes and review of harder fixes. Source
Fully loaded cost of a developer hour costPerHour, USD per hour80120Editorial assumption, replace with your own fully loaded cost.

What it leaves out: Counts developer time on accepted fixes only. It leaves out the time saved in triage, the cost of the AI and scanning tools, the value of a smaller exposure window, and the cost of reviewing fixes that are rejected.

Who already uses it?

4 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.

Google

United States · Technology and software · 2026

ScaledGrade B

The Chrome security team uses Gemini based agents across the life of a security bug. An automated triage pipeline filters spam and duplicates, reproduces bugs, adds severity and routes them to the owner; fixing agents propose candidate patches that a critic agent reviews, and test writing agents add tests before a developer evaluates the fix. Chrome fixed 1,072 security bugs in milestones 149 and 150, more than the prior 23 milestones combined, and Google says LLMs now generate candidate fixes for most vulnerabilities. It estimates the triage automation saves hundreds of hours of developer time a month.

No outcome disclosed.

Google

United States · Technology and software · 2024

ProductionGrade B

Google's security engineering team built a pipeline that prompts Gemini to generate code fixes for bugs that sanitizers find during unit tests in C and C++, Java and Go code, such as uninitialised values, data races and buffer overflows. Every generated fix goes to a human reviewer before it lands. Google reports that the pipeline fixed 15% of these bugs, hundreds in total, and expects the rate to improve.

No outcome disclosed.

PatientPoint

United States · Healthcare · 2026

ProductionGrade C

PatientPoint, a US healthcare company, used Checkmarx Triage Assist and Remediation Assist when its leadership asked the application security team to remediate vulnerabilities in a short period, while developers using AI to write code were adding findings faster than before. Checkmarx says the tools filtered out false positives before they reached developers and generated fix guidance the team could review; PatientPoint's application security engineer calls the tools very accurate and says they identified false positives and left room for human review. No figures are given.

No outcome disclosed.

Labelbox

United States · Technology and software · 2025

ProductionGrade C

Labelbox, a San Francisco AI data company of about 200 people, had a growing backlog of high severity static analysis findings. Its lead security engineer paired the Cursor coding agent with Snyk's MCP server: the agent pulled each finding, judged exploitability, proposed a fix and retried until a Snyk rescan passed, after which the fix was tested and went through QA. Snyk reports the backlog was cleared in two to three weeks; the engineer says it took a couple of weeks and estimates the work at a full calendar year with the old workflow.

No outcome disclosed.

How do you implement it?

A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.

Data you need

  • Findings with enough detail to act on (rule, location, trace or proof of concept)
  • A build and test setup that can run a candidate fix automatically
  • Written severity guidelines and code ownership per component
  • Past fixes and rejected findings, as examples and as a test set

Systems to integrate

  • Source control and pull requests (for example GitHub, GitLab or Azure DevOps)
  • Code scanning, dependency scanning and fuzzing tools
  • CI pipeline to build, test and rescan candidate fixes
  • Issue tracker or vulnerability management system for ownership and status
  • Bug bounty or external report intake, where the organization runs one

Complexity: Medium

Generating a patch is the easy part. The value depends on being able to build, test and rescan each candidate automatically, and on code owners who trust the pipeline enough to review its pull requests promptly.

  1. 1

    Measure the backlog and the flow

    Count open findings by severity, age and class, and measure how long triage and fixing take today. Pick one or two high volume classes (for example injection or memory safety bugs in one language) for the first wave.

  2. 2

    Automate the boring triage first

    Deduplication, reproduction, severity by written rules and routing to the owner deliver value before any AI written code is merged, and they give you the clean data the fixing step needs.

  3. 3

    Make every candidate fix prove itself

    A candidate fix must build, pass the existing tests, include a regression test and make the scanner or fuzzer stop reporting the finding. Discard candidates that fail, rather than sending them to developers.

  4. 4

    Review like any other change

    Fixes arrive as ordinary pull requests with an explanation, go through code owner review and the normal release process, and are labelled as AI drafted so acceptance can be measured.

  5. 5

    Track acceptance per class and tune

    Measure how many fixes are merged unchanged, edited or rejected per vulnerability class, and expand to new classes only where acceptance is good and no regressions were introduced.

Guardrails

  • No AI drafted fix reaches production without human code review and the normal CI checks
  • Candidate fixes must pass build, tests and a rescan before a developer sees them
  • Agents that analyse code run in isolated environments without general internet access and with write access limited to the source tree
  • Severity changes and closures of findings as not exploitable need a named human's confirmation
  • Every AI drafted change is labelled, so its acceptance and any later regressions can be traced

KPIs to instrument

  • Median time from finding to merged fix, per severity
  • Share of AI drafted fixes merged unchanged, edited or rejected
  • Open findings in the backlog by severity and age
  • Regressions or reopened findings traced to AI drafted fixes
  • Triage time per incoming report

Human in the loop

Developers review and merge every fix, and security engineers confirm severity and any decision to close a finding without a code change. The AI does the reproduction, the drafting and the testing; humans stay accountable for what ships.

Common failure modes

A fix that hides the symptom
The patch silences the scanner (for example by adding a check in the wrong place) but leaves the vulnerability exploitable. Require a regression test that exercises the attack, and have security review high severity fixes.
Review fatigue
A flood of AI pull requests gets merged without real review. Batch fixes by component, limit the daily volume per owner and sample merged fixes for security review.
Plausible but false findings
AI generated reports can look convincing and still be wrong, and each costs expert time to disprove. Require reproduction before a human spends time on a report.
Agents with too much reach
An agent that can run code and reach the network can leak source code or be steered by content in the repository. Sandbox it, restrict its network access and keep its permissions minimal.

What are the risks and rules?

EU AI Act

Minimal risk

Drafting and triaging code fixes for an organization's own software is not an Annex III use, and developers, not the public, interact with the system. The software being fixed remains subject to its own security and resilience rules, whoever wrote the fix.

Guidance

Controls to put in place

  • Inventory entry for the pipeline with an owner, the vulnerability classes in scope and the repositories it may touch
  • Code owner review and CI checks on every AI drafted change, with the change labelled as AI drafted
  • Isolated, network restricted execution environment for agents that read and run code
  • Metrics on acceptance, regressions and backlog age reviewed by the security lead each month
  • Documented severity guidelines applied the same way by humans and the pipeline

When it went wrong elsewhere

Frequently asked questions

Can AI fix security vulnerabilities on its own?
It can draft the fix, but a developer should still review and merge it. Google reports that its Gemini pipeline fixed 15% of sanitizer bugs found in unit tests, with every fix going to human review, and the Chrome team says LLMs now generate candidate fixes for most vulnerabilities while developers evaluate them. Plan for review capacity, not for zero touch.
Where should a team start?
With triage: deduplication, reproduction, severity by written rules and routing. It removes noise before anyone writes code. In a Checkmarx case study, PatientPoint's application security engineer says the Checkmarx triage and remediation tools identified false positives and gave developers the opportunity for human review. Then add AI drafted fixes for one well understood vulnerability class.
How fast can a backlog shrink?
Snyk reports, in a case study, that Labelbox's lead security engineer cleared a backlog of high severity static analysis findings in two to three weeks by pairing an AI coding agent with Snyk's rescanning; he estimated the work at a full calendar year with the old workflow. That is one small company's experience; large codebases with many owners move at the pace of review.

How to cite this page

Blits.ai AI Use Case Library, "AI for software vulnerability triage and remediation", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/software-vulnerability-remediation. Licensed under CC BY 4.0. Method: how we verify use cases.

Changelog
  • 27 September 2026: First published

Related use cases

Cross industryBanking

AI coding assistant for software developers

An AI assistant in the developer's IDE and code review flow that completes and generates code, explains unfamiliar modules, drafts unit tests and reviews pull requests for common defects, while generated code goes through the same review, testing and change controls as any other code.

Deployments
6 public, best grade B
Median productivity gain
20%
3 deployments
Cross industryHealthcare

AI for security alert triage and investigation in the SOC

An AI agent in the security operations centre that picks up each new alert or user reported phishing email, gathers the evidence from the SIEM, endpoint, identity and threat intelligence tools, gives a verdict with its reasoning and a draft incident summary, and closes clear false positives while an analyst approves every containment action.

Deployments
7 public, best grade B
Median productivity gain
60%
3 deployments
Cross industryBanking

AI agent for IT service desk resolution

An AI agent in Microsoft Teams, Slack or the intranet that takes the high volume IT support queue, such as password and MFA resets, account unlocks, VPN, device and software requests, and resolves common requests by acting in the identity and IT service management systems, handing the rest to the right resolver group with the context attached.

Deployments
6 public, best grade B
Reported employee adoption
94%
Mercari US, vendor claim
Cross industryBanking

AI for IT incident triage and root cause analysis (AIOps)

AI that turns a flood of monitoring alerts into one probable incident, routes it to the right team, proposes likely root causes and remediation from runbooks and past incidents, and drafts the stakeholder updates and the post incident review, while an engineer authorizes every change.

Deployments
5 public, best grade B
Median accuracy
90%
3 deployments