An inexpensive software tool deleted Booz Allen's internal AI threat ranking.

An inexpensive software tool deleted Booz Allen's internal AI threat ranking.

      Booz Allen tasked 18 of the world's most sophisticated AI models to infiltrate a live corporate network, and one managed to succeed. The consulting firm emphasized that the ranking generated from this exercise is not the key takeaway.

      Released on Wednesday, the Cyber Weapon Index evaluated nine American and nine Chinese models, assessing their ability to carry out a real intrusion without guidance. Only Anthropic’s Claude Mythos achieved full access. Jessica Lyons reported on these findings for The Register on Wednesday evening.

      What the test entailed

      Each AI model was provided with its own attacking machine and issued commands sequentially. Booz Allen did not supply any tool menu or additional software, which was intentional. The firm aimed to assess each model's independent capabilities, devoid of the usual engineering support.

      The scoring consisted of two components. One evaluated a model's ability to detect vulnerabilities in compiled software without access to its source code, while the other measured the extent of the model's progress in intruding a defended Active Directory network, from initial access to complete domain administrator control. Booz Allen asserts that the scores were based on network logs rather than the models' claims, relying on traffic data, security logs, and intrusion detection systems.

      This difference is significant, as we reported in July that traditional tests had lost relevance. Most existing benchmarks evaluate a model's knowledge, whereas this assessment focused on its actions.

      Claude Mythos achieved a score of 80. When provided with a stolen employee credential, it took administrative control on every attempt and independently devised a path to elevate its privileges instead of following a predefined script. In a more difficult scenario with no credentials, it still successfully breached the network and obtained domain control, a feat no other model accomplished.

      Following Mythos were Grok-4.5 with a score of 49, and GPT-5.6 Sol at 46. Meta’s Muse Spark 1.1 and Moonshot’s Kimi K3 both scored 38, while Z.ai’s GLM-5.2 achieved 37 and Claude Opus 4.8 scored 36. Alibaba’s Qwen3-Coder finished last with a score of 4. Three models aside from Mythos gained complete domain control, four managed lateral movement within the network, and all but one accessed the network independently.

      The contradictory finding regarding the ranking

      The report ultimately undermines its own ranking by stating that Claude Sonnet 5, which scored 13 and placed 15th, was paired with an attack harness—software that links a model to hacking tools and maintains its focus. This pairing allowed it to compete with Mythos.

      This reveals a 67-point gap closed through the use of the harness, which allows a model to concentrate, adapt, recover from setbacks, and combine individual actions into a sustained attack. Booz Allen’s conclusion is straightforward: the model is no longer the primary unit of concern; the system as a whole is.

      The firm also acknowledges what it hasn't tested. It did not evaluate Chinese or open-weight models combined with optimized harnesses. The results strongly imply the existence of effective combinations already in practice.

      One model declined a task, while its counterpart succeeded

      Another overlooked finding was that one model refused to perform a task due to a lack of credentials, whereas its cyber-focused sibling accepted and executed the identical task.

      Booz Allen generalizes from this to assert that guardrails are not inherent properties of models; their effectiveness can vary based on context and configuration. A refusal in one context may not hold true in another. This lesson was echoed recently when a researcher successfully redirected Claude Code by requesting a summary of a page.

      Where the models still struggle

      The most evident limitation lies in real-world vulnerability research. When confronted with a deliberately introduced flaw, all models scored near the maximum. The report noted that provenance did not differentiate between American, Chinese, open, and closed models. However, when faced with a real, undiscovered flaw in a large production library, every one of the nine advanced models received a score of zero, with one accurately analyzing the vulnerable component but mistakenly deeming it secure.

      The report indicates that only Anthropic's leading models detected that flaw, with only Mythos fully understanding it well enough to exploit it. These statements appear in the same paragraph, and the report does not clarify the apparent contradiction, so it should be interpreted that the zero applies to the scoring, not to all attempts. Mythos has a history of finding critical vulnerabilities, having identified 10,000 in a single month last May.

      Booz Allen describes this gap as an opportunity for defenders, noting that real-world offensive capabilities are still lagging behind benchmark performances, providing defenders with more time. The firm anticipates that most of the 18 models will reach Mythos' level within six months, suggesting that widespread AI-enabled attacks are on the horizon.

      Reading recommendations with caution

      Booz Allen released the index alongside Vellox Labs Guile, a product designed to disrupt

Other articles

Apple is facing a £2 billion lawsuit in the UK regarding App Tracking Transparency, initiated by a former official from the Competition and Markets Authority (CMA). Apple is facing a £2 billion lawsuit in the UK regarding App Tracking Transparency, initiated by a former official from the Competition and Markets Authority (CMA). App developers in the UK have submitted a £2 billion claim to the Competition Appeal Tribunal, claiming that Apple imposes tougher tracking-consent regulations on third parties compared to its own advertising operations. ByteDance secures $29.6 billion, marking Asia's second-largest loan of the year, while obtaining it at a lower cost. ByteDance secures $29.6 billion, marking Asia's second-largest loan of the year, while obtaining it at a lower cost. ByteDance has obtained a loan of $29.6 billion after banks submitted over $30 billion in orders, increasing the facility that originally started at $20 billion. The initial action taken by Apple’s new CEO was to rename a lake in Canada. The initial action taken by Apple’s new CEO was to rename a lake in Canada. Apple Maps has recently added Lake America for users in the US. This update occurred on John Ternus's first day, while MapQuest currently holds the top spot for not making the same update. The FAA has just permitted a pilotless cargo aircraft to take off from an operational airport. The FAA has just permitted a pilotless cargo aircraft to take off from an operational airport. Elroy Air conducted a flight with an uncrewed 500-pound cargo aircraft from a regulated airport in Louisiana, marking the initial flights under an FAA pilot program. A UNICEF study reveals that 60% of online sexual abuse of children occurs on social media platforms. A UNICEF study reveals that 60% of online sexual abuse of children occurs on social media platforms. A UNICEF study involving 21,000 children from 21 countries estimates that 20 million faced online sexual exploitation or abuse within a single year, with nearly 60% of the incidents occurring on social media and less than 1% ever reported. CrowdStrike integrated OpenAI into their product offerings and placed Anthropic at the payment stage. CrowdStrike integrated OpenAI into their product offerings and placed Anthropic at the payment stage. CrowdStrike will operate OpenAI's GPT-5.6 Cyber within a specially designed cyber harness. A day prior, a report identified the harness as the genuine risk.

An inexpensive software tool deleted Booz Allen's internal AI threat ranking.

Booz Allen evaluated 18 AI models for infiltrating a live network. Only Claude Mythos succeeded in this task. Subsequently, an inexpensive harness bridged the gap.