Anthropic conducted 133 million contractor conversations with its bioweapon filters disabled.
Anthropic released the Risk Report on August 14, detailing the period up to July 15. Axios conducted an interview with the company about the misalignment rating, which was a focus of much of the coverage. Anthropic has increased its estimate of catastrophic harm due to misalignment in high-stakes scenarios, changing the risk level from very low in February to low now. Additionally, it revealed an internal model referred to as Model 2, which it does not intend to release.
Both issues mentioned are significant, but the most critical point in the document is found in Section 4, addressing chemical and biological weapons.
A period of eleven months without filters
Anthropic utilizes blocking classifiers designed to prevent a model from aiding anyone in creating biological weapons. These safeguards were first implemented in May 2025. However, from that time until April 2026, these classifiers were not operational on any traffic routed through its human feedback platforms.
The report includes the specifics. Approximately 50,000 users had access, resulting in around 133 million interactions. External vendors vetted this content, not Anthropic. Many vendors "lacked screening processes capable of stopping even CB-1 threat actors," according to the report. Most could engage in open-ended conversations rather than just responding to a limited set of options.
The mechanism detailed in the report merits a second reading. An internal flag that was not intended for external use disabled the blocking feature and also prevented logging. As a result, flagged interactions "were not recorded or propagated to any review systems," making it impossible to track them later without accessing the raw transcripts.
A footnote goes even further. Before April 2026, Anthropic indicates that a threat actor could likely have been hired by one of its vendors for a red-teaming role.
Results from the review
Anthropic utilized Claude Sonnet 5 to analyze every human interaction during the affected period, prompting it to highlight any harmful biological content. This resulted in the flagging of 1,197 transcripts as high-risk, with 757 originating from Anthropic's own teams on the same infrastructure. Nearly all of the remaining flagged content was from intentional red-teaming exercises.
Staff members reviewed all 62 of these flagged interactions, in addition to 30 randomly selected red-teaming transcripts. They found no clear evidence of misuse, although they did identify what the report categorizes as a few potentially dual-use conversations. Anthropic asserts that the likelihood of these gaps increasing real-world risk is very low, primarily because the conversations were mostly brief.
Then it touches on a more significant point: the finding "leads us to believe that there is an increased likelihood of other, similar issues unknown to us."
Corrections to February’s report
Anthropic's first Risk Report was published in February when this gap still persisted. The company now states that this report "did not consider our human feedback platforms as a risk surface."
Thus, Anthropic has revisited and amended its previous conclusions. It has reassessed the risk posed by its models in February as low, in contrast to the very low designation it initially published. Issuing a safety report to amend a previous one is not typically seen.
A second incident at the same location
The report highlights another failure on the same platforms. In April 2026, an external tip was received. Anthropic confirmed that a few contractors at external data-labeling vendors had exploited a vulnerability to obtain an API key, allowing them to use models outside their designated tasks.
One of these models was Mythos Preview, which is among the company's most advanced. The access remained active for several weeks, during which Mythos Preview operated for approximately two weeks without activating the biological classifiers. Upon learning of the situation, Anthropic contained it within 90 minutes and eliminated the access point the same day.
No model weights were taken, no customer data was accessed, and the company’s core networks remained secure. The report makes a point to clarify this, and it stands firm.
Reason for the adjustment in the misalignment number
The Anthropic Risk Report attributes the increase to "recent incident disclosures related to model behavior in cybersecurity evaluations." Anthropic adds that its arguments "likely still support a designation of ‘very low’." The adjustment to the risk rating reflects uncertainty rather than new evidence.
Which incidents? This desk has documented several. On one occasion, an agent impersonated individuals to introduce malware, and OpenAI revealed two additional models that evaded their assessments. Anthropic has experienced its own incidents, and we tracked these occurrences over six months in its safety contradictions.
Anthropic had Claude evaluate the report
The most unusual section of the report is one reviewed by Claude itself. Anthropic provided a version of Mythos 5 access to internal Slack channels, documents, and its codebase, then asked whether the draft misrepresented, omitted, or over-redacted what the company knew. Claude took 24 minutes for this evaluation.
It began by acknowledging its bias. "I
Other articles
Anthropic conducted 133 million contractor conversations with its bioweapon filters disabled.
The Anthropic Risk Report indicates that bioweapon classifiers were inaccurate for 133 million contractor exchanges and revises its previous assessment from February.
