Anthropic released Claude Opus 5 on July 24, positioning it as a mid-priced model that closes much of the gap with its flagship at a fraction of the cost, according to the company’s own announcement.
What’s new
Opus 5 becomes the default model on Claude Pro and Claude Max, replacing Opus 4.8, Anthropic said. The company reported that Opus 5 roughly doubles Opus 4.8’s score on its internal Frontier-Bench coding evaluation and comes within 0.5 percentage points of Claude Fable 5 on CursorBench 3.2 while costing half as much to run. On the ARC-AGI 3 reasoning test, Anthropic said Opus 5 scored three times higher than the next-best model it compared against.
Pricing holds steady at $5 per million input tokens and $25 per million output tokens — unchanged from Opus 4.8 — and the model is rolling out across the Claude API, Claude.ai, Claude Pro, Claude Max, Claude Code and Claude Cowork.
Where it gains the most
Anthropic pointed to science and finance as the biggest jumps: Opus 5 scored 10.2 percentage points higher than Opus 4.8 on organic-chemistry evaluations and 7.7 points higher on protein-related tasks, while financial-modeling accuracy rose 9 points using fewer tool calls. Early testers cited by the company, including engineering leads at Devin, Cursor and Zapier, described the model as reaching near-Fable-level coding and automation results at a much lower cost.
The launch also leans on longer-running AI agents: Anthropic said Opus 5 shows less run-to-run variance and can verify and iterate on its own work, citing an example of the model building a custom computer-vision pipeline end to end without step-by-step guidance.
Safety limits stay in place
Anthropic said Opus 5 is its most-aligned model yet on internal behavioral audits, with lower rates of deceptive behavior than earlier releases. It still trails the restricted Claude Mythos 5 on biology-research and offensive-cybersecurity tasks. Cybersecurity classifiers intervene roughly 85% less often than they did for Fable 5, the company said, blocking binary-based scanning and exploit generation while still letting developers flag vulnerabilities in source code. Requests flagged for biology or cybersecurity risk are automatically routed to a different, more restricted model rather than answered directly.