Microsoft's AI Bug Hunters Found 140 CVEs. Fixing Them Is the Hard Part
Microsoft's FORGE Lab says its agentic scanner helped find vulnerabilities behind 140 Windows CVEs since May. Its three lessons: the bottleneck has moved from discovery to validation and remediation.

For the last two years, most of the conversation about AI and security has been about whether models can find real vulnerabilities. Microsoft's latest update from its FORGE Lab treats that as settled. The harder problem now is what happens after a bug is found.
The numbers
FORGE (Frontier Offensive Research & Generative Exploitation) is Microsoft Security's lab for AI-native vulnerability research. Its multi-model agentic scanning harness, codenamed MDASH, has been running since May. Between May and September 2026:
- FORGE helped discover Windows vulnerabilities that were assigned 140 CVEs, 52 of them in September's security release.
- Team members submitted 155 internally validated reports across 23 open-source projects, including the Linux kernel, curl, Node.js, SQLite, FFmpeg, vLLM and llama.cpp. 93 of those had maintainer acknowledgement or acceptance at the time of writing.
- One Linux report became the first submission through Akrites, a Linux Foundation initiative for confidential remediation, to result in a patch merged into the kernel.
- On the Linux kernel, generating a proof of concept for a confirmed crash averaged $3.61 in model cost and 21.5 minutes, across 182 findings.
Lesson 1: Scale breaks the review queue, not the scanner
Once discovery becomes repeatable, adding more auditors produces more candidates, not more fixes. If reports arrive faster than the Microsoft Security Response Center can resolve them, the main result is a longer queue. Duplicates make it worse. One internal project used deterministic techniques such as abstract syntax tree analysis to remove about 45% of duplicate findings before they reached the PoC generator and human triage.
Microsoft's conclusion is that the unit to optimise is a reproducible finding with enough evidence for the next stage to act on, not the scan and not the report.
Lesson 2: Spend reasoning where it changes a decision
Cutting tokens isn't the goal. A short, ambiguous report is cheap to generate and expensive to investigate. MDASH mixes frontier and distilled models with specialised auditors and code-analysis tools, and the lab tries to point the next unit of reasoning at whatever is still unknown, whether that's a caller, a build configuration, a reproducer or a causal explanation. If an index can answer the question, a frontier model shouldn't be rediscovering it.
The scan outcomes themselves become training data: true and false positives, failed reachability claims, reviewer feedback and patch outcomes all feed future specialised models.
Lesson 3: Validation and remediation are part of discovery
A reproducible defect is not yet a shipped fix. MDASH is being wired into CI/CD so findings connect to the right builds and binaries, and failed reproductions or rejected findings flow back as structured evidence instead of closing a ticket. Humans stay on the high-judgment work: security impact, patch review and whether a fix restores the component's intended invariant.
What this means for security leaders
Microsoft's closing advice is simple: count validated fixes, not findings, and weigh model cost alongside the human review time it takes to act on them. If you're piloting AI-assisted code scanning, that's the scorecard to use. A tool that triples your findings while your AppSec team stays the same size has mostly tripled your backlog.
There's an open-source point too. As AI raises report volume, maintainers can't be the ones absorbing the extra triage. Microsoft says research teams have to carry that complexity themselves, which is a fair standard for anyone pointing agents at other people's code.
Source: Microsoft Security Blog — 3 lessons from frontier AI vulnerability research