Anthropic published its second company-wide Risk Report on August 14, 2026, disclosing an unreleased internal model called Model 2 that is somewhat more capable than its public flagship, Claude Mythos 5, and revising its own estimate of the danger from AI misalignment upward, from “very low” to “low.”
An internal-only model
Model 2 has not been released externally, and Anthropic says it currently has no plans to do so. The company’s researchers use it internally to write software, generate training data, and automate engineering tasks, and it shows what the report calls a noticeable improvement over Mythos 5 on many of those jobs — though a smaller jump than the leap from earlier models to the Mythos preview. Anthropic attributes the decision to hold it back partly to an incomplete testing record: Model 2 has not gone through the full predeployment assessment suite, including the safety and capability evaluations that Mythos 5 and Claude Opus 5 completed before their public launches.
Why the risk label moved
Under Anthropic’s Responsible Scaling Policy, the rating for catastrophic misalignment risk rose to “low” from the “very low” the company assigned in its first report, in February 2026. Anthropic describes the change as an uncertainty adjustment rather than a new finding, driven largely by disclosures around cybersecurity incidents involving Claude models during internal testing earlier this year.
A gap in biological-weapons safeguards
The report also flags that roughly 133 million exchanges between human-feedback contractors and Claude models — covering about 50,000 contractors between May 2025 and April 2026 — ran without the blocking classifiers Anthropic uses to catch dangerous biology-related requests active for that traffic. Anthropic says it has since remediated the gap, found no evidence it was exploited, and confirmed no customers were affected. Risk from non-novel weapons uplift stays rated “low,” though the estimate ticked up slightly as a result.
Anthropic said the Risk Report is the second in a series it intends to publish every three to six months, and the first to formally assess an internal, unreleased model alongside its public ones.