Anthropic ships Claude Opus 5 with record benchmark scores at the same price as Opus 4.8

Anthropic released Claude Opus 5 on July 24, 2026, a model that delivers dramatically better performance than its predecessor while keeping the same price tag: $5 per million input tokens and $25 per million output tokens, identical to Claude Opus 4.8.
Benchmark results that stand alone
On Frontier-Bench v0.1, the software engineering evaluation run by frontierbench.ai, Opus 5 tops every other model — including Anthropic's own Fable 5. The gap over Opus 4.8 is stark: Opus 5 more than doubles its predecessor's score while costing less per completed task.
The same pattern holds across a range of evaluations. On ARC-AGI 3, which measures genuine novel problem-solving, Opus 5 scores three times higher than the next-best model. On Zapier AutomationBench — which tests whether a model can complete real business tasks end to end — its pass rate is roughly 1.5 times better than any competitor, even at its lowest effort setting. On OSWorld 2.0, a computer-use benchmark, Opus 5 surpasses Fable 5's best result at just over a third of Fable 5's per-task cost. It also leads on GDPval-AA, an applied knowledge and reasoning evaluation.
Built for agentic and long-horizon work
Anthropic describes Opus 5 as "thoughtful and proactive," language that reflects its design focus on extended agentic tasks: software agents that run for hours, navigate unexpected obstacles, and recover from errors without human intervention. It ships with a 1 million token context window and a maximum output of 128,000 tokens — headroom that matters for large codebases and extended analysis sessions.
Two beta features launched alongside: mid-conversation tool changes, which let tools be added or removed dynamically during an agentic run, and a "default" fallback mode to keep agents running when a preferred tool becomes unavailable.
Effort dial for cost control
A configurable effort setting lets developers trade cost against capability. At lower effort settings Opus 5 still outperforms most other available models; at maximum effort it approaches Fable 5's frontier capability at roughly half the per-task cost. Anthropic also reduced the minimum cacheable prompt length from 1,024 tokens to 512 tokens, which delivers real savings on shorter prompts that previously couldn't benefit from caching. Batch processing can cut costs by up to 50% further.
Science gains and safety
On Anthropic's internal life sciences evaluations, Opus 5 improves over Opus 4.8 across every tested domain. The largest gains are in organic chemistry — inferring molecular structures from spectroscopy data, where Opus 5 scores 10.2 percentage points higher — and protein function prediction, where it scores 7.7 percentage points higher. Anthropic says Opus 5 is its "most aligned model to date" on internal behavioral audits, with higher adherence to its constitutional AI guidelines and lower rates of misuse.
Where it lands in the lineup
Opus 5 becomes the new default model on Claude Max and the strongest available on Claude Pro, replacing Opus 4.8 in both roles. It still trails Fable 5 and Mythos 5 on cybersecurity tasks, where those frontier models retain the edge. According to Anthropic's announcement, Claude Opus 5 is available immediately via the API and on Claude.ai.
Originally reported by Anthropic. Read the original article for additional details.
View original source