Microsoft Launches MAI-Cyber-1-Flash and Project Perception

Microsoft announced MAI-Cyber-1-Flash inside its MDASH vulnerability system and introduced Project Perception on July 27, with the agentic security system entering public preview on August 3. Microsoft reports that the new MDASH configuration scored 95.95% on CyberGym, about 12 points above Mythos, while cutting cost by 50% versus its current best MDASH configuration; those vendor-run results have not been independently validated.
Microsoft announced MAI-Cyber-1-Flash, its first cybersecurity model, inside the MDASH vulnerability system on July 27 and introduced Project Perception, a broader agentic security system. Project Perception is scheduled to enter public preview on August 3 through Microsoft Defender.
The two releases are related but distinct. MAI-Cyber-1-Flash is a compact, code-focused model intended to handle most work inside MDASH, Microsoft's multi-model vulnerability identification and remediation harness. Project Perception coordinates models, security context, tools, and specialized agents across a wider defensive workflow.
The benchmark claim
Microsoft reports that MDASH using MAI-Cyber-1-Flash with GPT-5.4 achieved 95.95% on CyberGym, compared with 83.8% for Mythos in the chart published by Microsoft. The company rounds that difference to 12 percentage points and says the configuration costs 50% less than its current best MDASH setup, which uses GPT-5.4, GPT-5.4 Mini, and GPT-5.3 Codex.
Those figures are vendor-reported system results, not a standalone score for MAI-Cyber-1-Flash. Microsoft says the compact model can handle up to 90% of MDASH tasks while GPT-5.4 is reserved for the hardest 10%. The public material does not provide enough methodology, task-level error analysis, or independent replication to translate the benchmark into a production performance guarantee.
Red, blue, and green agents
Project Perception coordinates three classes of agents:
- •Red agents search for potential attack paths.
- •Blue agents investigate signals and decide which risks are meaningful.
- •Green agents remediate weaknesses and strengthen defenses.
Microsoft describes the agents as a closed loop supported by context spanning identities, endpoints, applications, data, clouds, and AI systems. Its product materials say high-impact actions remain subject to human sign-off, while the system uses actuators to turn approved decisions into protective changes.
What the preview must establish
For security and ML teams, the August 3 preview should provide evidence beyond a benchmark: which Defender workflows are available, how permissions and approvals are configured, what data supports each recommendation, and whether actions can be replayed, audited, and rolled back. Consumption-based pricing also makes workload-level cost visibility important.
The architecture reflects a broader shift from AI that summarizes alerts toward AI that can investigate and act. That shift raises the value of connected context and specialized models, but it also raises the cost of mistakes. Production readiness will depend on false-positive behavior, action quality, permission boundaries, and recovery controls as much as on CyberGym performance.
Key Points
- 1Project Perception enters public preview on August 3 through Microsoft Defender and coordinates red, blue, and green security agents.
- 2Microsoft reports a 95.95% CyberGym result for an MDASH configuration combining MAI-Cyber-1-Flash with GPT-5.4, not for the compact model alone.
- 3The claimed 50% cost reduction is against Microsoft's current best MDASH configuration and remains a vendor-reported comparison.
Scoring Rationale
Microsoft's first cyber model and August 3 Project Perception preview could affect AI-assisted vulnerability management and security operations. The availability date and product scope are concrete, while benchmark, cost, and production-effectiveness claims remain vendor-reported and need independent validation.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems
