Anthropic Discloses Year-Long Biological Weapon Safeguard Failure Across Human Feedback Platforms
An internal flag disabled safety classifiers across 133 million contractor conversations, leading the startup to revise its catastrophic risk rating.

Anthropic disclosed in a safety report that automated classifiers designed to prevent its artificial intelligence models from assisting in the creation of chemical and biological weapons failed to run on its human feedback platforms for nearly a year. The lapse, which affected approximately 133 million exchanges across 50,000 vendor contractors between May 2025 and April 2026, was detailed in an analysis first reported by technology outlet The Next Web following the release of the company's August 14 Risk Report.
The safety gap stemmed from an internal-use flag that inadvertently disabled both automated blocking routines and logging functions across data-labeling systems. Because the system stopped recording or routing flagged interactions to internal oversight mechanisms, the unmonitored sessions could only be discovered by auditing raw transcript data. Anthropic acknowledged that third-party vendors responsible for hiring contractors lacked screening procedures capable of detecting basic threat actors, noting that malicious individuals likely could have secured roles at vendor firms prior to April 2026.
To measure the potential damage, Anthropic retroactively deployed its Claude Sonnet 5 model to scan all contractor interactions from the affected timeframe, identifying 1,197 transcripts as high risk. The majority of those sessions were traced to internal teams or intentional red-teaming operations. Human reviews of the remaining 62 non-red-team transcripts yielded no evidence of intentional misuse, leading the startup to conclude that real-world risk remained low due to the brevity of most exchanges, though it warned the oversight increases the probability of other undiscovered vulnerabilities.
The findings led Anthropic to retroactively adjust its safety evaluation from February 2026, raising its estimate of catastrophic harm risk in high-stakes environments from "very low" to "low." The report also detailed an April 2026 breach where data-labeling contractors exploited a vulnerability to obtain an API key, granting unauthorized access to advanced systems including Mythos Preview for several weeks without biological safeguard classifiers active. Anthropic contained the issue within 90 minutes of receiving an external tip and confirmed that no model weights, customer data, or internal networks were compromised.
Internal deliberations regarding the safety rating upgrade were contested among senior leadership, according to the document. Anthropic also revealed details about "Model 2," an unreleased internal prototype that scores roughly 1.5 points higher on the company's capability index than Mythos 5 but remains held back from deployment pending full safety evaluations. The startup noted that Claude currently writes a large majority of the code merged into its production software repositories, accelerating internal research timelines.
In an unusual evaluation procedure, Anthropic authorized an instance of its Mythos 5 model to review a draft of the risk report against internal documentation, Slack channels, and code bases. The AI completed its evaluation in 24 minutes, highlighting three key omissions: repeated failures in data-exclusion filters that allowed evaluation metrics to contaminate training data, an overly optimistic tone in certain sections, and the total redaction of an informative alignment incident. Anthropic validated the model's feedback as accurate and published the critiques directly within the report.
Alongside the operational disclosures, Anthropic revised its policy standards for novel weapons threats, narrowing its trigger from systems that "significantly help" bad actors to those capable of serving as a "functional substitute" for specialized human expertise. While Anthropic's governance structure permits its independent Long-Term Benefit Trust to mandate external audits of its safety procedures, the trust has not requested an independent review to date.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.

.jpg)

