Claude launches Fable 5.1 and Mythos 5.1 with big benchmark gains

Claude has just rolled out Fable 5.1 and Mythos 5.1, positioning them for coding, research, cybersecurity, and life sciences. The company touts higher Terminal-Bench and OSWorld scores versus prior models, plus lower cache-read costs and new Enterprise Frontier Safeguards.

claude cover

TL;DR

  • New models: Claude Fable 5.1 (complex coding/research) and Claude Mythos 5.1 (cyberdefense, life science)
  • Benchmark gains (Fable 5.1): Terminal-Bench-Science 0.1 52.6%; Terminal-Bench 4.0 55.8%; OSWorld 2.0 77.9%/41.7%
  • Additional scores (Fable 5.1): GDPval-AA v2 1,853, AutomationBench 31.4%, CursorBench 3.2.0 73.4%, HLE 60.9%/65.0%
  • Mythos 5.1 benchmark: Terminal-Bench 4.0 60.9%
  • Evaluation caveats: production safeguards enabled; some tasks zeroed by interventions; Terminal-Bench-Science 0.1 standard error ±3.5–4.5
  • Cost and safety updates: 75% cheaper cache reads vs Fable 5; Enterprise Frontier Safeguards phased rollout starting fall 2026; fewer benign flags and fallbacks

Claude has introduced Claude Fable 5.1 and Claude Mythos 5.1, positioning them as models for coding and knowledge work. The company claims Fable 5.1 is suited to complex, long-running tasks and research, while Mythos 5.1 targets cyberdefenders and life scientists.

Claude reports that Fable 5.1 scored 52.6% on Terminal-Bench-Science 0.1, compared with 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol. On Terminal-Bench 4.0, Fable 5.1 reached 55.8%, ahead of Fable 5 at 42.0%, Opus 5 at 52.3% and GPT-5.6 Sol at 37.3%.

Mythos 5.1 scored 60.9% on Terminal-Bench 4.0, according to Claude’s benchmark data. Fable 5.1 also posted higher scores than its predecessor on GDPval-AA v2, AutomationBench and CursorBench 3.2.0:

  • GDPval-AA v2: 1,853 for Fable 5.1, compared with 1,723 for Fable 5 and 1,824 for Opus 5
  • AutomationBench: 31.4% for Fable 5.1, compared with 17.1% for Fable 5, 26.9% for Opus 5 and 19.6% for GPT-5.6 Sol
  • CursorBench 3.2.0: 73.4% for Fable 5.1, compared with 70.5% for Fable 5, 70.0% for Opus 5 and 67.2% for GPT-5.6 Sol

On OSWorld 2.0, Fable 5.1 scored 77.9% in the partial setting and 41.7% in the strict setting. Fable 5 scored 72.9% and 36.1%, respectively, while Opus 5 scored 75.4% and 39.6%.

Fable 5.1 also exceeded the earlier model on Humanity’s Last Exam, scoring 60.9% without tools and 65.0% with tools. Fable 5 scored 57.8% and 63.8%, while Opus 5 scored 56.6% and 63.6%.

The results come with qualifications. Claude notes that Fable 5.1 was evaluated with production safeguards enabled, and that safeguards intervened on some OSWorld 2.0 tasks, producing zero scores for Fable 5.1 and Fable 5. Safeguards also affected some AutomationBench tasks, resulting in a zero score for Fable 5. Cybersecurity tasks affected by safeguards were completed by Claude Opus 4.8, while biology tasks were completed by Claude Opus 5. The Terminal-Bench-Science 0.1 results carry a reported standard error of ±3.5–4.5 points per model.

Claude is also highlighting lower operating costs. Cache reads for Fable 5.1 cost 75% less than those for Fable 5, which the company estimates reduces total model costs by around 25% for typical workloads and by as much as 45% for highly agentic workloads. Its cost charts compare benchmark scores with mean cost per task across low, medium, high, extra-high and maximum effort settings.

At lower effort levels, Claude’s charts indicate that Fable 5.1 can approach or exceed Fable 5’s results at lower cost. On Terminal-Bench-Science 0.1, Fable 5.1’s scores range from approximately 26% at low effort to 52.5% at maximum effort, compared with roughly 12% to 25% for Fable 5. On Terminal-Bench 4.0, Mythos 5.1 ranges from approximately 42% to 61%, while Fable 5.1 ranges from about 40% to 55.5%.

The release also introduces Enterprise Frontier Safeguards, which Claude describes as providing enterprise customers with privacy equivalent to zero data retention while retaining protections against adversarial use. The safeguards are scheduled to roll out in phases beginning in fall 2026.

Claude additionally claims that its cybersecurity safeguards now flag benign requests about 60% less often. The fallback rate for basic biology and medical questions has reportedly declined by around 85%.

Fable 5.1 is available everywhere, while Mythos 5.1 is available through trusted access programs for cyberdefenders and life scientists.

Source: Claude on X

Continue the conversation on Slack

Did this article spark your interest? Join our community of experts and enthusiasts to dive deeper, ask questions, and share your ideas.

Join our community