What problem does it solve?
Security tools find far more problems than teams can fix. Static analysis (SAST), dependency scanners, fuzzers, penetration tests and bug bounty reports all feed the same backlog, and each finding needs someone to confirm it is real, work out how serious it is, find the owner and write a fix that does not break anything. Developers see security tickets as interruptions to feature work, so the backlog grows and old findings stay open.
AI is making the imbalance sharper on both sides. Models are now used to find vulnerabilities (Google's Big Sleep agent found bugs in Chrome's V8 engine in 2025). Separately, the Chrome security team reports receiving more bug reports by March 2026 than in all of 2025, without naming a single cause, and says triaging a single report used to take 5 to 30 or more minutes of expert time. At the same time, low quality AI generated reports waste maintainers' time, and developers using coding assistants ship more code, and so more findings, than before. Fixing has to scale as fast as finding, without letting unreviewed code into production.
- The Chrome security team reports that historically, triaging a single security report took anywhere from 5 to 30 or more minutes and relied primarily on human expertise.Stronger with every update: How we're making Chrome and the web safer in the AI Era (2026)
How does it work?
- Collect and deduplicate. Findings arrive from scanners, fuzzers, CI pipelines and external reports. The system drops spam and duplicates and checks that each one describes a real security issue in scope.
- Validate and rank. Where possible it reproduces the bug (running the proof of concept or the failing test), checks whether the vulnerable code is reachable, and assigns a severity using written guidelines. Findings that mitigations already neutralise are marked for a human to confirm and close.
- Route to the owner. The finding goes to the right component and code owner with the stack trace, severity and context attached.
- Draft the fix. A fixing agent proposes one or more candidate patches; a separate critic step or reviewer model checks them against the style guide, and the fix is rebuilt and rescanned to prove the finding is gone. Test writing agents add a regression test.
- Human review and merge. The developer reviews the pull request like any other change and merges, edits or rejects it. Rejections and edits feed back into the prompts and examples.
- Audience
- Employee facing
- Autonomy
- Copilot
- Adoption
- Early adopters
- Channels
- Internal tools, API and system to system
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
No public deployment has disclosed a measurable outcome yet.
Value drivers: Risk and loss reduction, Employee productivity, Speed and cycle time, Compliance quality.
Indicative value
A software organization that fixes 2,000 to 5,000 security findings a year
USD 12,000 to USD 336,000
Developer time released on security fixes, valued at cost per year
How this is calculated
Formula: findingsPerYear * hoursPerFix * aiFixShare * timeSavedShare * costPerHour. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Security findings fixed per year findingsPerYear, findings per year | 2,000 | 5,000 | Editorial assumption for a mid sized engineering organization. Replace with the count from your vulnerability management or code scanning tool. |
| Developer hours per manual fix, including review hoursPerFix, hours per finding | 1.5 | 4 | GitHub reports a median of 1.5 hours to resolve code scanning alerts manually, measured in its public beta on alerts raised in pull requests on new code; the high end is an editorial allowance for older backlog findings, which take longer. Source |
| Share of findings where an AI drafted fix is accepted aiFixShare, fraction of findings | 0.1 | 0.2 | An editorial range around the only measured acceptance rate, Google's 15% of sanitizer bugs fixed by its LLM pipeline with human review. Chrome's candidate fixes for most vulnerabilities are drafts, not accepted fixes, so they are not used here. Source |
| Share of developer time saved on those findings timeSavedShare, fraction of fix time | 0.5 | 0.7 | GitHub reports a median of 28 minutes with Copilot Autofix against 1.5 hours manually, about 69% less, measured in its public beta on new alerts in pull requests; the low end allows for backlog fixes and review of harder fixes. Source |
| Fully loaded cost of a developer hour costPerHour, USD per hour | 80 | 120 | Editorial assumption, replace with your own fully loaded cost. |
What it leaves out: Counts developer time on accepted fixes only. It leaves out the time saved in triage, the cost of the AI and scanning tools, the value of a smaller exposure window, and the cost of reviewing fixes that are rejected.
Who already uses it?
4 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
United States · Technology and software · 2026
The Chrome security team uses Gemini based agents across the life of a security bug. An automated triage pipeline filters spam and duplicates, reproduces bugs, adds severity and routes them to the owner; fixing agents propose candidate patches that a critic agent reviews, and test writing agents add tests before a developer evaluates the fix. Chrome fixed 1,072 security bugs in milestones 149 and 150, more than the prior 23 milestones combined, and Google says LLMs now generate candidate fixes for most vulnerabilities. It estimates the triage automation saves hundreds of hours of developer time a month.
No outcome disclosed.
United States · Technology and software · 2024
Google's security engineering team built a pipeline that prompts Gemini to generate code fixes for bugs that sanitizers find during unit tests in C and C++, Java and Go code, such as uninitialised values, data races and buffer overflows. Every generated fix goes to a human reviewer before it lands. Google reports that the pipeline fixed 15% of these bugs, hundreds in total, and expects the rate to improve.
No outcome disclosed.
PatientPoint
United States · Healthcare · 2026
PatientPoint, a US healthcare company, used Checkmarx Triage Assist and Remediation Assist when its leadership asked the application security team to remediate vulnerabilities in a short period, while developers using AI to write code were adding findings faster than before. Checkmarx says the tools filtered out false positives before they reached developers and generated fix guidance the team could review; PatientPoint's application security engineer calls the tools very accurate and says they identified false positives and left room for human review. No figures are given.
No outcome disclosed.
Labelbox
United States · Technology and software · 2025
Labelbox, a San Francisco AI data company of about 200 people, had a growing backlog of high severity static analysis findings. Its lead security engineer paired the Cursor coding agent with Snyk's MCP server: the agent pulled each finding, judged exploitability, proposed a fix and retried until a Snyk rescan passed, after which the fix was tested and went through QA. Snyk reports the backlog was cleared in two to three weeks; the engineer says it took a couple of weeks and estimates the work at a full calendar year with the old workflow.
No outcome disclosed.
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- Findings with enough detail to act on (rule, location, trace or proof of concept)
- A build and test setup that can run a candidate fix automatically
- Written severity guidelines and code ownership per component
- Past fixes and rejected findings, as examples and as a test set
Systems to integrate
- Source control and pull requests (for example GitHub, GitLab or Azure DevOps)
- Code scanning, dependency scanning and fuzzing tools
- CI pipeline to build, test and rescan candidate fixes
- Issue tracker or vulnerability management system for ownership and status
- Bug bounty or external report intake, where the organization runs one
Complexity: Medium
Generating a patch is the easy part. The value depends on being able to build, test and rescan each candidate automatically, and on code owners who trust the pipeline enough to review its pull requests promptly.
- 1
Measure the backlog and the flow
Count open findings by severity, age and class, and measure how long triage and fixing take today. Pick one or two high volume classes (for example injection or memory safety bugs in one language) for the first wave.
- 2
Automate the boring triage first
Deduplication, reproduction, severity by written rules and routing to the owner deliver value before any AI written code is merged, and they give you the clean data the fixing step needs.
- 3
Make every candidate fix prove itself
A candidate fix must build, pass the existing tests, include a regression test and make the scanner or fuzzer stop reporting the finding. Discard candidates that fail, rather than sending them to developers.
- 4
Review like any other change
Fixes arrive as ordinary pull requests with an explanation, go through code owner review and the normal release process, and are labelled as AI drafted so acceptance can be measured.
- 5
Track acceptance per class and tune
Measure how many fixes are merged unchanged, edited or rejected per vulnerability class, and expand to new classes only where acceptance is good and no regressions were introduced.
Guardrails
- No AI drafted fix reaches production without human code review and the normal CI checks
- Candidate fixes must pass build, tests and a rescan before a developer sees them
- Agents that analyse code run in isolated environments without general internet access and with write access limited to the source tree
- Severity changes and closures of findings as not exploitable need a named human's confirmation
- Every AI drafted change is labelled, so its acceptance and any later regressions can be traced
KPIs to instrument
- Median time from finding to merged fix, per severity
- Share of AI drafted fixes merged unchanged, edited or rejected
- Open findings in the backlog by severity and age
- Regressions or reopened findings traced to AI drafted fixes
- Triage time per incoming report
Human in the loop
Developers review and merge every fix, and security engineers confirm severity and any decision to close a finding without a code change. The AI does the reproduction, the drafting and the testing; humans stay accountable for what ships.
Common failure modes
- A fix that hides the symptom
- The patch silences the scanner (for example by adding a check in the wrong place) but leaves the vulnerability exploitable. Require a regression test that exercises the attack, and have security review high severity fixes.
- Review fatigue
- A flood of AI pull requests gets merged without real review. Batch fixes by component, limit the daily volume per owner and sample merged fixes for security review.
- Plausible but false findings
- AI generated reports can look convincing and still be wrong, and each costs expert time to disprove. Require reproduction before a human spends time on a report.
- Agents with too much reach
- An agent that can run code and reach the network can leak source code or be steered by content in the repository. Sandbox it, restrict its network access and keep its permissions minimal.
What are the risks and rules?
EU AI Act
Minimal risk
Drafting and triaging code fixes for an organization's own software is not an Annex III use, and developers, not the public, interact with the system. The software being fixed remains subject to its own security and resilience rules, whoever wrote the fix.
Guidance
- NIST SP 800-218, Secure Software Development Framework (SSDF) Version 1.1 (NIST, North America). Recommended practices for mitigating software vulnerabilities, including identifying, analysing and remediating them; the process an AI remediation pipeline has to fit into.
- NIST SP 800-218A, Secure Software Development Practices for Generative AI and Dual-Use Foundation Models (NIST, North America). An SSDF community profile that adds practices for developing generative AI systems, useful when the remediation pipeline itself is built on models.
- Secure by Design (Cybersecurity and Infrastructure Security Agency (CISA), North America). CISA's programme asking every technology provider to take ownership of product security at the executive level, with alerts on eliminating specific vulnerability types such as cross site scripting and OS command injection.
Controls to put in place
- Inventory entry for the pipeline with an owner, the vulnerability classes in scope and the repositories it may touch
- Code owner review and CI checks on every AI drafted change, with the change labelled as AI drafted
- Isolated, network restricted execution environment for agents that read and run code
- Metrics on acceptance, regressions and backlog age reviewed by the security lead each month
- Documented severity guidelines applied the same way by humans and the pipeline
When it went wrong elsewhere
- The I in LLM stands for intelligence (curl project on AI generated vulnerability reports). The curl maintainer describes AI generated bug bounty reports that looked plausible but were hallucinated, and how each one takes a human's time to disprove; the reason to require reproduction before triage.
Frequently asked questions
- Can AI fix security vulnerabilities on its own?
- It can draft the fix, but a developer should still review and merge it. Google reports that its Gemini pipeline fixed 15% of sanitizer bugs found in unit tests, with every fix going to human review, and the Chrome team says LLMs now generate candidate fixes for most vulnerabilities while developers evaluate them. Plan for review capacity, not for zero touch.
- Where should a team start?
- With triage: deduplication, reproduction, severity by written rules and routing. It removes noise before anyone writes code. In a Checkmarx case study, PatientPoint's application security engineer says the Checkmarx triage and remediation tools identified false positives and gave developers the opportunity for human review. Then add AI drafted fixes for one well understood vulnerability class.
- How fast can a backlog shrink?
- Snyk reports, in a case study, that Labelbox's lead security engineer cleared a backlog of high severity static analysis findings in two to three weeks by pairing an AI coding agent with Snyk's rescanning; he estimated the work at a full calendar year with the old workflow. That is one small company's experience; large codebases with many owners move at the pace of review.
How to cite this page
Blits.ai AI Use Case Library, "AI for software vulnerability triage and remediation", last verified 27 September 2026, https://www.blits.ai/ai-use-cases/software-vulnerability-remediation. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 27 September 2026: First published