Skip to content Skip to footer

How to Turn AI Vulnerability Scanning into Verified Security Fixes

What Happened

Microsoft Security’s FORGE Lab reports that its multi-model scanning harness, MDASH, helped discover Windows vulnerabilities assigned 140 CVEs from May through September 2026, including 52 addressed in September’s security release. It also submitted 155 internally validated reports across 23 open-source projects; 93 received documented maintainer acknowledgement or acceptance. These are distinct measures: a validated report is not necessarily an accepted report or a released fix. [2]

The research highlights a cost gap. In a Linux kernel study, successful proof-of-concept generation for 182 confirmed crashes averaged $3.61 in model cost and 21.5 minutes. Those figures exclude initial screening, unsuccessful candidates, and human work. Separately, Unit 42 reports that attackers are using Web3 infrastructure in campaigns targeting enterprise cloud environments and exploiting the open-source supply chain; the available account does not establish a single attack technique or prevalence rate. [1] [2]

Why It Matters to Businesses

AI can increase the volume of plausible vulnerability findings faster than teams can reproduce, triage, patch, and release them. A dashboard counting findings may therefore overstate security progress. The more useful business measure is validated fixes deployed, alongside reviewer time and the rate of false or duplicate reports. FORGE reduced duplicate findings across scans by about 45% in one internal project, illustrating the value of improving the workflow around scanning. [2]

Cloud risk also extends beyond application code: open-source dependencies and infrastructure used in attacks must be considered together when setting detection and response priorities. [1]

Kimbodo Engineering Perspective

We would use AI to investigate bounded questions—not to replace security ownership. A model is useful when asked to test whether a suspected flaw is reachable, produce a reproducible case, or check whether a patch closes that case. Broad, repeated scans are less valuable when they generate duplicates or findings that cannot reach a maintainer or release pipeline. Human review remains necessary for severity, exploitability, patch safety, and disclosure decisions. [2]

How We Would Implement It

  • Inventory critical repositories, dependencies, cloud assets, and owners; select scanning targets by exposure and business impact.
  • Run scanners and AI-assisted analysis in isolated, least-privilege environments. Require each finding to include the affected version, reproduction steps, evidence, and a proposed owner.
  • Deduplicate findings against prior reports and route uncertain cases to focused tests or specialist review rather than another broad scan.
  • Track each issue through reproduction, review, patch, regression test, and deployment. Feed reviewer verdicts and CI results back into scanning rules, while keeping release approval with accountable engineers. [2]
  • Monitor dependency changes and cloud activity as part of the same response workflow, without assuming every unusual Web3-related indicator is malicious. [1]

Risks, Costs and Security

The main cost is not just model inference; it is failed investigations, reviewer capacity, and the time required to ship a safe fix. AI-generated proof-of-concept code and repository contents should be treated as untrusted: restrict credentials and network access, log tool actions, and require approval before code changes or external disclosure. Measure outcomes by fixes verified in production, not by the number of findings an agent can produce. [2]

Where Kimbodo Comes In

Kimbodo builds and operates this in production for businesses — see our AI Security & Guardrails practice, or Request a Security Review.

Sources

  1. [1] Evolution of Web3 in Cloud Supply Chain Attacks
  2. [2] 3 lessons from frontier AI vulnerability research

Leave a comment

0.0/5