Microsoft developed a security system featuring AI agents from red, blue, and green teams. It will be available for public preview starting August 3.
TL;DR: Microsoft has introduced Project Perception, an agent-based security system featuring red, blue, and green AI agents. Its MAI-Cyber-1-Flash model achieves a 96% score on CyberGym—12 points higher than Mythos—while costing 50% less. A public preview is available starting August 3.
On Monday, Microsoft unveiled Project Perception, an agent-based security system that coordinates three types of AI agents in a continuous loop: red team agents identify vulnerabilities before attackers can, blue team agents evaluate and determine the significance of risks, and green team agents implement defensive fixes across the environment. The system will enter public preview on August 3. Microsoft characterized this as a "new Cyber Stack" designed for an era where AI-driven attacks progress quicker than human responses.
The initial performance benchmark showcases Microsoft's MAI-Cyber-1-Flash model, integrated into its MDASH software vulnerability management system, which has achieved a 96% score on CyberGym—a recognized standard for vulnerability assessment. This is a 12-point advantage over Anthropic's Mythos, the leading cybersecurity frontier model. Additionally, Microsoft asserts that this configuration provides approximately 50% savings compared to the existing MDASH production setup by aligning specific models to certain tasks instead of funneling all tasks through a single, costly frontier model.
Project Perception employs a multi-model architecture rather than one model for all functions. Frontier models tackle complex reasoning, while specialized cyber models focus on high-volume and low-latency tasks. The system leverages Microsoft's extensive visibility across identities, endpoints, applications, data, clouds, and AI systems, allowing it to take action in those environments instead of merely generating alerts. Microsoft's AI has already detected a record number of vulnerabilities in its software, and Project Perception expands this capability to customer environments through continuously operating agents instead of relying on monthly patch cycles.
The competitive jab at Anthropic is intentional. Mythos gained prominence as the standard cybersecurity model after being temporarily obstructed by the White House, attracting considerable attention for its offensive abilities. Microsoft claims its specialized model outperforms Mythos at half the cost, utilizing training data from years of defending enterprise environments that Anthropic lacks. This month, the White House launched Gold Eagle to coordinate AI-powered cyber defense, and Project Perception represents Microsoft's effort to become the foundational platform for defense operations. The public preview on August 3 will allow enterprises to test it, raising the question of whether the 96% benchmark can withstand real-world attacks as opposed to synthetic assessments.
Other articles
Microsoft developed a security system featuring AI agents from red, blue, and green teams. It will be available for public preview starting August 3.
Project Perception employs AI agents that continuously target, examine, and resolve security vulnerabilities. Microsoft's MAI-Cyber-1-Flash achieves a score of 96% on CyberGym, outperforming Mythos by 12 points.
