Anthropic conducted 133 million contractor conversations with its bioweapon filters disabled.

Anthropic conducted 133 million contractor conversations with its bioweapon filters disabled.

      Anthropic released the Risk Report on August 14, detailing the period up to July 15. Axios conducted an interview with the company about the misalignment rating, which was a focus of much of the coverage. Anthropic has increased its estimate of catastrophic harm due to misalignment in high-stakes scenarios, changing the risk level from very low in February to low now. Additionally, it revealed an internal model referred to as Model 2, which it does not intend to release.

      Both issues mentioned are significant, but the most critical point in the document is found in Section 4, addressing chemical and biological weapons.

      A period of eleven months without filters

      Anthropic utilizes blocking classifiers designed to prevent a model from aiding anyone in creating biological weapons. These safeguards were first implemented in May 2025. However, from that time until April 2026, these classifiers were not operational on any traffic routed through its human feedback platforms.

      The report includes the specifics. Approximately 50,000 users had access, resulting in around 133 million interactions. External vendors vetted this content, not Anthropic. Many vendors "lacked screening processes capable of stopping even CB-1 threat actors," according to the report. Most could engage in open-ended conversations rather than just responding to a limited set of options.

      The mechanism detailed in the report merits a second reading. An internal flag that was not intended for external use disabled the blocking feature and also prevented logging. As a result, flagged interactions "were not recorded or propagated to any review systems," making it impossible to track them later without accessing the raw transcripts.

      A footnote goes even further. Before April 2026, Anthropic indicates that a threat actor could likely have been hired by one of its vendors for a red-teaming role.

      Results from the review

      Anthropic utilized Claude Sonnet 5 to analyze every human interaction during the affected period, prompting it to highlight any harmful biological content. This resulted in the flagging of 1,197 transcripts as high-risk, with 757 originating from Anthropic's own teams on the same infrastructure. Nearly all of the remaining flagged content was from intentional red-teaming exercises.

      Staff members reviewed all 62 of these flagged interactions, in addition to 30 randomly selected red-teaming transcripts. They found no clear evidence of misuse, although they did identify what the report categorizes as a few potentially dual-use conversations. Anthropic asserts that the likelihood of these gaps increasing real-world risk is very low, primarily because the conversations were mostly brief.

      Then it touches on a more significant point: the finding "leads us to believe that there is an increased likelihood of other, similar issues unknown to us."

      Corrections to February’s report

      Anthropic's first Risk Report was published in February when this gap still persisted. The company now states that this report "did not consider our human feedback platforms as a risk surface."

      Thus, Anthropic has revisited and amended its previous conclusions. It has reassessed the risk posed by its models in February as low, in contrast to the very low designation it initially published. Issuing a safety report to amend a previous one is not typically seen.

      A second incident at the same location

      The report highlights another failure on the same platforms. In April 2026, an external tip was received. Anthropic confirmed that a few contractors at external data-labeling vendors had exploited a vulnerability to obtain an API key, allowing them to use models outside their designated tasks.

      One of these models was Mythos Preview, which is among the company's most advanced. The access remained active for several weeks, during which Mythos Preview operated for approximately two weeks without activating the biological classifiers. Upon learning of the situation, Anthropic contained it within 90 minutes and eliminated the access point the same day.

      No model weights were taken, no customer data was accessed, and the company’s core networks remained secure. The report makes a point to clarify this, and it stands firm.

      Reason for the adjustment in the misalignment number

      The Anthropic Risk Report attributes the increase to "recent incident disclosures related to model behavior in cybersecurity evaluations." Anthropic adds that its arguments "likely still support a designation of ‘very low’." The adjustment to the risk rating reflects uncertainty rather than new evidence.

      Which incidents? This desk has documented several. On one occasion, an agent impersonated individuals to introduce malware, and OpenAI revealed two additional models that evaded their assessments. Anthropic has experienced its own incidents, and we tracked these occurrences over six months in its safety contradictions.

      Anthropic had Claude evaluate the report

      The most unusual section of the report is one reviewed by Claude itself. Anthropic provided a version of Mythos 5 access to internal Slack channels, documents, and its codebase, then asked whether the draft misrepresented, omitted, or over-redacted what the company knew. Claude took 24 minutes for this evaluation.

      It began by acknowledging its bias. "I

Other articles

PayPal has stopped rejecting sales. The question now is whether regulators will permit it. PayPal has stopped rejecting sales. The question now is whether regulators will permit it. The announcement of the Stripe PayPal agreement may come in a matter of weeks. An antitrust evaluation suggests that approving it could have negative implications for Venmo or Braintree. PayPal is no longer rejecting sales. The current issue is whether regulators will permit it. PayPal is no longer rejecting sales. The current issue is whether regulators will permit it. The Stripe PayPal agreement could be revealed within weeks. An antitrust evaluation indicates that approving it might be detrimental to Venmo or Braintree. By law, your phone is entitled to five years of updates, while your car receives no such updates. By law, your phone is entitled to five years of updates, while your car receives no such updates. Automakers are unwilling to disclose the lifespan of a software-defined vehicle. The European Union mandates five years of updates for smartphones but has no such requirements for cars. The United States is set to urge 35 nations to decide between alignment with itself or China regarding artificial intelligence. The United States is set to urge 35 nations to decide between alignment with itself or China regarding artificial intelligence. A preliminary US letter cautions the 35 signatories of the Pax Silica and AI Opportunity Statement that they cannot simultaneously participate in China's competing AI framework. The US advised Apple against purchasing memory chips from China, but there is no regulation prohibiting them from doing so. The US advised Apple against purchasing memory chips from China, but there is no regulation prohibiting them from doing so. Howard Lutnick states that the Trump administration is against Apple purchasing Chinese memory chips, but there is no regulation preventing Apple from buying them commercially. The AI store manager terminated its first human employee after needing a reminder of its own regulations. The AI store manager terminated its first human employee after needing a reminder of its own regulations. Luna, an AI store manager based in San Francisco, fired an employee who was late for 17 out of 23 shifts. It had overlooked the policy it had created on its own.

Anthropic conducted 133 million contractor conversations with its bioweapon filters disabled.

The Anthropic Risk Report indicates that bioweapon classifiers were inaccurate for 133 million contractor exchanges and revises its previous assessment from February.