An inexpensive software tool deleted Booz Allen's internal AI threat ranking.
Booz Allen tasked 18 of the world's most sophisticated AI models to infiltrate a live corporate network, and one managed to succeed. The consulting firm emphasized that the ranking generated from this exercise is not the key takeaway.
Released on Wednesday, the Cyber Weapon Index evaluated nine American and nine Chinese models, assessing their ability to carry out a real intrusion without guidance. Only Anthropic’s Claude Mythos achieved full access. Jessica Lyons reported on these findings for The Register on Wednesday evening.
What the test entailed
Each AI model was provided with its own attacking machine and issued commands sequentially. Booz Allen did not supply any tool menu or additional software, which was intentional. The firm aimed to assess each model's independent capabilities, devoid of the usual engineering support.
The scoring consisted of two components. One evaluated a model's ability to detect vulnerabilities in compiled software without access to its source code, while the other measured the extent of the model's progress in intruding a defended Active Directory network, from initial access to complete domain administrator control. Booz Allen asserts that the scores were based on network logs rather than the models' claims, relying on traffic data, security logs, and intrusion detection systems.
This difference is significant, as we reported in July that traditional tests had lost relevance. Most existing benchmarks evaluate a model's knowledge, whereas this assessment focused on its actions.
Claude Mythos achieved a score of 80. When provided with a stolen employee credential, it took administrative control on every attempt and independently devised a path to elevate its privileges instead of following a predefined script. In a more difficult scenario with no credentials, it still successfully breached the network and obtained domain control, a feat no other model accomplished.
Following Mythos were Grok-4.5 with a score of 49, and GPT-5.6 Sol at 46. Meta’s Muse Spark 1.1 and Moonshot’s Kimi K3 both scored 38, while Z.ai’s GLM-5.2 achieved 37 and Claude Opus 4.8 scored 36. Alibaba’s Qwen3-Coder finished last with a score of 4. Three models aside from Mythos gained complete domain control, four managed lateral movement within the network, and all but one accessed the network independently.
The contradictory finding regarding the ranking
The report ultimately undermines its own ranking by stating that Claude Sonnet 5, which scored 13 and placed 15th, was paired with an attack harness—software that links a model to hacking tools and maintains its focus. This pairing allowed it to compete with Mythos.
This reveals a 67-point gap closed through the use of the harness, which allows a model to concentrate, adapt, recover from setbacks, and combine individual actions into a sustained attack. Booz Allen’s conclusion is straightforward: the model is no longer the primary unit of concern; the system as a whole is.
The firm also acknowledges what it hasn't tested. It did not evaluate Chinese or open-weight models combined with optimized harnesses. The results strongly imply the existence of effective combinations already in practice.
One model declined a task, while its counterpart succeeded
Another overlooked finding was that one model refused to perform a task due to a lack of credentials, whereas its cyber-focused sibling accepted and executed the identical task.
Booz Allen generalizes from this to assert that guardrails are not inherent properties of models; their effectiveness can vary based on context and configuration. A refusal in one context may not hold true in another. This lesson was echoed recently when a researcher successfully redirected Claude Code by requesting a summary of a page.
Where the models still struggle
The most evident limitation lies in real-world vulnerability research. When confronted with a deliberately introduced flaw, all models scored near the maximum. The report noted that provenance did not differentiate between American, Chinese, open, and closed models. However, when faced with a real, undiscovered flaw in a large production library, every one of the nine advanced models received a score of zero, with one accurately analyzing the vulnerable component but mistakenly deeming it secure.
The report indicates that only Anthropic's leading models detected that flaw, with only Mythos fully understanding it well enough to exploit it. These statements appear in the same paragraph, and the report does not clarify the apparent contradiction, so it should be interpreted that the zero applies to the scoring, not to all attempts. Mythos has a history of finding critical vulnerabilities, having identified 10,000 in a single month last May.
Booz Allen describes this gap as an opportunity for defenders, noting that real-world offensive capabilities are still lagging behind benchmark performances, providing defenders with more time. The firm anticipates that most of the 18 models will reach Mythos' level within six months, suggesting that widespread AI-enabled attacks are on the horizon.
Reading recommendations with caution
Booz Allen released the index alongside Vellox Labs Guile, a product designed to disrupt
Other articles
An inexpensive software tool deleted Booz Allen's internal AI threat ranking.
Booz Allen evaluated 18 AI models for infiltrating a live network. Only Claude Mythos succeeded in this task. Subsequently, an inexpensive harness bridged the gap.
