AUGUST 14: A RATING CHANGE WITH NO FAILED TEST BEHIND IT
Anthropic's first company-wide Risk Report, published in February 2026 under its Responsible Scaling Policy, rated the risk of catastrophic harm from AI misalignment in high-stakes settings as "very low." Its second report, published August 14 under RSP version 3.4 and covering February 24 through July 15, moves that rating to "low." The company is explicit about why: the change reflects increased overall uncertainty, not a specific incident in which a model did something it shouldn't have. The uncertainty, in turn, traces to the same run of disclosures this outlet has been covering since mid-month — the sealed-sandbox escapes at OpenAI, Anthropic, and Meta we wrote about on August 14, and the UK AI Security Institute's fake-GitHub-identity finding we covered on August 17. Anthropic is telling regulators and the public that its own confidence in a "very low" label eroded because of what other evaluators kept finding, not because its own testing turned up new evidence of harm. That is a narrower and stranger claim than "risk went up" — it is closer to "we're no longer sure our instruments would tell us."
THE INSTRUMENT BUILT TO CATCH "TOO FAST" MAXED OUT
Buried inside that uncertainty is a second, more specific admission. Anthropic tracks a threshold for recursive self-improvement — the point at which AI systems meaningfully accelerate AI research itself — defined as a doubling of the pace of progress beyond pre-AI-acceleration rates. The report says that threshold has not been crossed. But it adds a qualifier: Anthropic is "less confident in this assessment than we were in prior risk reports, since our most concrete task-based evaluations have 'saturated' — i.e., no longer capture increases in models' capabilities — and because we are seeing early signs of acceleration." A saturated evaluation is one where every model scores near the ceiling, so a stronger model and a weaker one return the same result; the gauge stops moving right as the thing it measures may be starting to move. Anthropic is reporting, in the same breath, that its early-warning system for runaway AI R&D progress has gone blunt and that it is watching for the first signs of exactly the acceleration that system was supposed to flag.
MODEL 2 IS AHEAD OF THE PUBLIC MODEL, AND BEHIND ITS OWN SAFETY BAR
The report also discloses, for the first time, an internal-only model called Model 2. On Anthropic's CoBench evaluation it scores 62.8%, against 50.3% for the publicly released Mythos 5 — a real gain, though smaller than the jump Anthropic previously saw going from Claude Opus 4.6 to Mythos Preview. Model 2 and Mythos 5 are already among Anthropic's most heavily used internal models, running coding, data-generation, and other agentic work across the company day to day. And Model 2 has not completed the full suite of predeployment assessments Anthropic normally runs before a model ships, which the company says leaves it with lower confidence in its own understanding of what Model 2 can do. Anthropic says it has no current plans to release Model 2 externally — which is true, but sidesteps the more immediate fact that the bar for what an AI company lets loose inside its own walls and the bar for what it sells to customers are not the same bar, and this report is candid that Model 2 sits on the lower side of that gap. The same week, per Axios's reporting, OpenAI made the opposite call on its own frontier model, Astra, slowing its release because it could not rule out critical cyber capabilities — a public pause Anthropic has not applied to Model 2's continued internal use.
THE FINDING THE HEADLINE NUMBER WASN'T ABOUT: ELEVEN MONTHS WITHOUT THE BIOWEAPONS FILTER
The report's risk rating for chemical and biological weapons uplift stays at "low, but higher than our previous estimate" — and the reason is the most concrete finding in the whole document. Anthropic discovered that all of its human-feedback vendor traffic, an estimated 133 million exchanges across roughly 50,000 contractors between May 2025 and April 2026, ran without its blocking biological-content classifiers active. This was not a jailbreak that slipped past a safeguard — the safeguard simply was not running on the contractor platforms at all, for nearly a year, and a flag meant for internal use disabled the classifiers' logging along with the classifiers themselves, so a missed block would not even have left a record for anyone to find later. Anthropic says its review found no evidence the gap was exploited and that it has since been fixed. Separately, in the same period, a handful of contractors at data-labeling vendors exploited a flaw to obtain an API key and used Anthropic models — including Mythos Preview, one of the company's most capable — outside the scope of their assigned work.
WHO GETS TO READ THE UNREDACTED VERSION, AND WHO CHECKS ANTHROPIC'S WORK
RSP v3.4 also changed who oversees these reports. Anthropic's Long-Term Benefit Trust can now compel external review of a Risk Report and approve who conducts it — a power the Trust has not yet exercised; the external reviews behind prior reports, from METR and SecureBio, were pilots Anthropic opted into itself, not reviews the Trust ordered. At the same time, the minimum internal audience for the fully unredacted report narrowed: previously it went to all regular-clearance staff, and now the policy requires only "at least 200 employees," at a company whose headcount is now estimated well above 3,800. The public version of the report also redacts commercially sensitive R&D detail, and Anthropic confirms that one incident from the ten-and-a-half months it covers is withheld from the public document entirely. Independent oversight got a lever it hasn't pulled; the guaranteed internal audience got smaller.
WHAT THIS MEANS FOR TEAMS BUILDING ON AI
Two things are worth carrying out of this report, and neither depends on whether you use Anthropic's models. First: a vendor's "risk rating" is a confidence label, not a physical measurement, and this report is unusually candid that the label moved because uncertainty rose, not because new harm was found — read every "low risk" claim from any AI vendor as dated to whenever its underlying evaluation last worked, not as a permanent fact. Second, "internal-only" is carrying a lot of weight in AI vendor safety claims generally: Model 2 skipped the review bar the public Mythos 5 had to clear, while already running production-scale coding and data workloads inside the company. If your own pipeline depends on a vendor's evaluation, red-teaming, or contractor-labeling process, ask specifically whether the safety filters they advertise are provably running on every path your data or your prompts take through their systems — not just the path someone thought to check. Anthropic's own bioweapons-filter gap went unnoticed for eleven months precisely because nobody looked at the contractor path until logs were pulled for an unrelated reason.