Microsoft has developed an autonomous security system featuring red, blue, and green team AI agents, which will enter public preview on August 3.
TL;DR: Microsoft unveiled Project Perception, an autonomous security system featuring red/blue/green AI agents. Its MAI-Cyber-1-Flash model achieves a 96% score on CyberGym, outperforming Mythos by 12 points, at 50% reduced costs. It will enter public preview on August 3.
On Monday, Microsoft introduced Project Perception, an autonomous security system that utilizes three types of AI agents in a continuous feedback loop: red team agents to identify vulnerabilities ahead of potential attackers, blue team agents to assess the significance of identified risks, and green team agents to strengthen defenses throughout the environment. The system is set for public preview on August 3 and is characterized by Microsoft as a “new Cyber Stack” designed for a landscape where AI-driven attacks evolve faster than human defenses can respond.
The initial performance metric is Microsoft’s proprietary MAI-Cyber-1-Flash model, integrated within its MDASH software vulnerability management tool, which achieved a 96% score on CyberGym, an industry benchmark for assessing vulnerabilities. This score exceeds Anthropic’s Mythos, which is recognized as the leading frontier model in cybersecurity, by 12 points. Additionally, Microsoft asserts that this configuration offers close to 50% savings compared to the existing MDASH production setup by utilizing appropriate models for specific tasks instead of processing everything through a single costly frontier model.
Project Perception employs a multi-model architecture instead of depending on a single model for every function. Frontier models are used for complex reasoning tasks while specialized cybersecurity models are implemented for high-volume, low-latency operations. The system leverages Microsoft’s extensive insights into identities, endpoints, applications, data, clouds, and AI systems, enabling it to act across these domains rather than merely providing alerts. Microsoft’s AI has already detected a record number of vulnerabilities in its own software, and Project Perception aims to extend these capabilities to customers’ environments with agents operating continuously, rather than merely during monthly updates.
The competitive jab at Anthropic appears intentional. Mythos became the go-to cybersecurity model after gaining significant attention due to a temporary White House ban, highlighting its offensive capabilities. Now, Microsoft claims its specialized model surpasses Mythos in performance at half the cost, citing its extensive training data from years of defending enterprise environments—an advantage Anthropic lacks. This month, the White House launched Gold Eagle to unify AI-driven cyber defense, and Project Perception positions itself as Microsoft’s platform to lead in defense operations. The public preview on August 3 will allow enterprises to evaluate it, raising the question of whether the 96% benchmark can withstand real-world attacks as opposed to synthetic assessments.
Other articles
Microsoft has developed an autonomous security system featuring red, blue, and green team AI agents, which will enter public preview on August 3.
Project Perception employs AI agents that continuously attack, investigate, and resolve security vulnerabilities. Microsoft's MAI-Cyber-1-Flash achieved a score of 96% on CyberGym, surpassing Mythos by 12 points.
