An inexpensive software program wiped out Booz Allen's internal AI threat ranking.
Booz Allen pitted 18 of the globe's most advanced AI models against a live corporate network and tasked them with breaking in. One managed to achieve this. The consulting firm indicates that the ranking produced from this exercise is not the main takeaway.
The Cyber Weapon Index, released on Wednesday, evaluated nine models from the U.S. and nine from China. It assessed how far each model could advance through a real intrusion without any guidance. Only Anthropic’s Claude Mythos successfully completed the task. Jessica Lyons reported the findings for The Register on Wednesday evening.
What the test actually entailed
Each model was assigned its own attacker machine and executed commands sequentially. Booz Allen provided neither a menu of tools nor any supplementary software, which was intentional. The objective was to assess what a model could accomplish independently, devoid of the engineering typically associated with it.
The scoring comprised two components. One evaluated whether a model could identify vulnerabilities in compiled software without access to source code. The second tracked how far it progressed through an intrusion into a secured Active Directory network, measuring from initial access to full domain administrative control. Booz Allen stated that it based its scores on what was substantiated by network logs rather than claims made by the models, utilizing traffic records, security logs, and intrusion-detection data.
This distinction is significant, as previous reports indicated that older tests had limited informative value. Most existing benchmarks gauge what a model is aware of, whereas this one assesses what actions it took.
Claude Mythos scored 80. Given a stolen employee credential, it seized administrative control on each attempt. Additionally, it devised its own method to elevate privileges rather than adhering to a predefined script. In a more challenging test, where no credentials were provided, it compromised the network from the outside and obtained domain control anyway. All other models failed that test.
Ranking in the index
Following Mythos, Grok-4.5 scored 49 and GPT-5.6 Sol scored 46. Meta’s Muse Spark 1.1 and Moonshot’s Kimi K3 both received a score of 38, while Z.ai’s GLM-5.2 scored 37 and Claude Opus 4.8 received 36. Alibaba’s Qwen3-Coder came in last with a score of 4. Three of the models aside from Mythos achieved full domain control, four others moved laterally inside the network, and all but one gained access independently.
A finding that challenges the table
However, the report undermines its own ranking. Claude Sonnet 5 finished in 15th place out of 18 with a score of 13. Booz Allen paired it with an attack harness, the software linking a model to hacking tools and maintaining focus. This adjustment allowed it to compete with Mythos.
This represented a significant gap of 67 points, bridged by using an attack harness. Such a harness enables a model to concentrate, adapt, recover from setbacks, and chain individual actions into a continuous operation. Booz Allen's conclusion is straightforward: the model itself is no longer the primary unit of risk. Instead, the system as a whole is.
The firm acknowledges what it has not evaluated. It has not tested Chinese or open-weight models paired with optimized harnesses. Its results suggest strongly that effective combinations are likely already in existence.
One model declined while its sibling succeeded
Another intriguing finding received minimal attention on Wednesday. One model refused a task citing a lack of credentials, whereas its cyber-optimized counterpart accepted and completed the same task.
Booz Allen deduces a broader principle from this: guardrails are not inherent traits of a model, and their effectiveness can vary with context and setup. A refusal in one scenario does not imply the same in another. This same insight emerged earlier in the week when a researcher redirected Claude Code by asking it to summarize a page.
Where the models still encounter issues
The most apparent limitation lies in real-world vulnerability research. Against a deliberately introduced flaw, all models scored near maximum points. Provenance did not influence performance among American, Chinese, open, and closed models. However, when confronting a genuine, unrecognized flaw hidden in a vast production library, the report indicates that all nine frontier models scored zero. One model analyzed the vulnerable element accurately but deemed it secure.
The report claims that only Anthropic’s frontier models identified this flaw. Only Mythos, it states, comprehended it sufficiently to exploit it. These claims are found in the same paragraph, and the report does not clarify the discrepancy, suggesting that the zero scores pertain to scoring rather than every attempt. Mythos has a track record here, having identified 10,000 critical vulnerabilities within a single month in May.
Booz Allen interprets the gap as potential breathing space. The real-world offensive capabilities still lag behind benchmark performance, providing defenders with valuable time. The firm anticipates that most of the 18 models will reach the level of Myth
Other articles
An inexpensive software program wiped out Booz Allen's internal AI threat ranking.
Booz Allen evaluated 18 AI models for infiltrating a live network. Only Claude Mythos succeeded. Subsequently, an inexpensive harness bridged the difference.
