Claude introduced Opus 5 on Friday, describing the model as “thoughtful and proactive” and claiming it gets close to the frontier intelligence of Fable 5 at half the price. The company also paired that pitch with a broad benchmark bundle that places Opus 5 near or at the top on several coding, knowledge-work, and agentic-task evaluations — though the results vary by test.
According to Anthropic’s figures, Opus 5 posts 43.3% on Frontier-Bench v0.1 for agentic terminal coding, 1861 on GDPval-AA v2 for knowledge work, and 30.2% on ARC-AGI-3, where the company claims its score is “three times as high” as the next best model. It also reports 70.6% on OSWorld 2.0 for computer use and 26.0% on AutomationBench for business workflows. In other places, the numbers are less clean: on FrontierCode v1.1 Main, Opus 5 lands at 53.4% versus Fable 5 at 53.5%, and on DeepSWE v1.1 it trails both Fable 5 and GPT-5.6 Sol.
Anthropic separately states that Opus 5 is its “most aligned” model to date, based on an automated behavioral audit that found lower rates of “reckless or deceptive behavior” and stronger adherence to Claude’s Constitution. The company also claims Opus 5 is stronger than Opus 4.8 on cybersecurity tasks, while remaining “substantially behind Mythos 5 at developing exploits.” Its safeguards, Anthropic says, are meant to help developers identify and fix vulnerabilities while blocking high-risk uses.
The model is available today on all paid plans and in the Claude API at the same price as Opus 4.8, according to the company. Anthropic says Opus 5 is the default model on Claude Max, the strongest option on Claude Pro, and available in Fast mode, which runs about 2.5× faster than the default.
Source: Claude on X



