Mandiant Discloses Agentic Vulnerability Discovery Harness Results
Google Threat Intelligence Group disclosed its internal Agentic Vulnerability Discovery Harness on August 18 after using it for 10 months in code-review and incident-response work. According to Google's technical post, the multi-agent system found more than 100 true-positive critical vulnerabilities in two days during an investigation involving stolen corporate repositories, and its broader use has resulted in 12 assigned CVEs.
Google Threat Intelligence Group has disclosed an internal framework called the Agentic Vulnerability Discovery Harness (AVDH), which Mandiant uses to analyze source code during proactive reviews, penetration tests, red-team operations, and incident-response engagements.
According to an August 18 Google Threat Intelligence post by Alex Tselevich and Michael Maturi, AVDH found more than 100 true-positive critical vulnerabilities in two days during a recent incident-response investigation involving stolen corporate repositories. Google reported that the results were achieved in a fraction of the time required for manual review.
Multi-agent review with human validation
Google describes AVDH as a multi-agent orchestration framework that structures source-code analysis, applies skeptical validation steps, and incorporates subject-matter expertise into the pipeline. The company characterizes the system as an augmentation tool for routine vulnerability discovery and validation, enabling human reviewers to focus their impact.
The published architecture is explicitly described as an internal, point-in-time design rather than a released product or an open-source framework. Google also states that AVDH can operate alongside CodeMender's ongoing scanning as a two-layer defense approach.
Reported scale and disclosures
Google reported that it has operated AVDH for 10 months across environments containing tens of millions of lines of code. During that period, the company said it executed thousands of pipelines and generated tens of thousands of findings.
According to the post, this activity identified dozens of assignable flaws in widely used web extensions and open-source projects, leading to 12 assigned CVEs and roughly a dozen additional vulnerabilities in active disclosure. The source does not identify the affected projects or CVE identifiers in the supplied material.
For security engineering teams, the reported results underscore a broader pattern in agentic application security: high-throughput code analysis is useful only when workflows include validation, prioritization, and accountable human review. Multi-agent systems can expand coverage across large repositories, but false positives, exploitability assessment, disclosure handling, and remediation still require disciplined security processes.
Google framed the disclosure against the growing risk that stolen proprietary code can be examined by attackers using automated AI tooling. The AVDH account provides a concrete example of defenders applying agent orchestration to reduce the time between repository exposure, vulnerability discovery, and validation.
Key Points
- 1Google reported that AVDH identified more than 100 true-positive critical vulnerabilities during a two-day investigation of stolen corporate repositories.
- 2The framework combines multi-agent orchestration with human subject-matter expertise and skeptical validation steps, rather than treating model output as final security judgment.
- 3Comparable agentic security workflows can increase repository coverage, but validation, disclosure coordination, and remediation remain essential operational controls.
Scoring Rationale
The disclosure offers a notable, operational example of agentic AI applied to large-scale source-code security review, with unusually strong reported vulnerability-discovery results. It is particularly relevant to application-security and ML engineering teams evaluating agent orchestration, although AVDH remains an internal architecture rather than a generally available tool.
Sources
Primary source and supporting public references used for this report.
Practice with real Ad Tech data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Ad Tech problems