Anthropic launches Claude Opus 5 with bold benchmark claims

Anthropic has just rolled out Claude Opus 5, touting near-frontier performance at half the price of Fable 5. The company’s benchmark bundle puts it near the top in coding, knowledge work, and agentic tasks—though results vary by test.

claude cover

TL;DR

  • Opus 5 launch: Positioned as “thoughtful and proactive”; claimed near Fable 5 frontier at half the price
  • Benchmarks (selected): Frontier-Bench v0.1 43.3%; GDPval-AA v2 1861; ARC-AGI-3 30.2%
  • Computer/workflows: OSWorld 2.0 70.6%; AutomationBench 26.0%
  • Mixed comparative results: FrontierCode v1.1 Main 53.4% vs Fable 5 53.5%; DeepSWE v1.1 trails Fable 5, GPT-5.6 Sol
  • Safety claims: “Most aligned” via automated behavioral audit; lower reckless/deceptive behavior; stronger Constitution adherence
  • Availability and modes: On paid plans and Claude API today; same price as Opus 4.8; Fast mode ~2.5× faster

Claude introduced Opus 5 on Friday, describing the model as “thoughtful and proactive” and claiming it gets close to the frontier intelligence of Fable 5 at half the price. The company also paired that pitch with a broad benchmark bundle that places Opus 5 near or at the top on several coding, knowledge-work, and agentic-task evaluations — though the results vary by test.

According to Anthropic’s figures, Opus 5 posts 43.3% on Frontier-Bench v0.1 for agentic terminal coding, 1861 on GDPval-AA v2 for knowledge work, and 30.2% on ARC-AGI-3, where the company claims its score is “three times as high” as the next best model. It also reports 70.6% on OSWorld 2.0 for computer use and 26.0% on AutomationBench for business workflows. In other places, the numbers are less clean: on FrontierCode v1.1 Main, Opus 5 lands at 53.4% versus Fable 5 at 53.5%, and on DeepSWE v1.1 it trails both Fable 5 and GPT-5.6 Sol.

Anthropic separately states that Opus 5 is its “most aligned” model to date, based on an automated behavioral audit that found lower rates of “reckless or deceptive behavior” and stronger adherence to Claude’s Constitution. The company also claims Opus 5 is stronger than Opus 4.8 on cybersecurity tasks, while remaining “substantially behind Mythos 5 at developing exploits.” Its safeguards, Anthropic says, are meant to help developers identify and fix vulnerabilities while blocking high-risk uses.

The model is available today on all paid plans and in the Claude API at the same price as Opus 4.8, according to the company. Anthropic says Opus 5 is the default model on Claude Max, the strongest option on Claude Pro, and available in Fast mode, which runs about 2.5× faster than the default.

Source: Claude on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community