Anthropic has introduced Claude Sonnet 5.5, the second model in the Claude 5.5 family. The company describes it as a faster, lower-cost successor to Sonnet 5, aimed at well-scoped tasks such as bug fixing, document creation, slide generation, spreadsheets, and interface design.
Claude claims Sonnet 5.5 generates output more than 30% faster than its predecessor, making it the fastest Sonnet model to date. Its listed price remains the same as Sonnet 5, but Anthropic states that the new model generally uses fewer tokens and costs up to 30% less per task.
The company also claims that Sonnet 5.5 can outperform Sonnet 5’s best scores on several benchmarks at Low or Medium effort for about one-tenth of the cost. Those comparisons should be read alongside Anthropic’s disclosures about testing issues, including a pre-release deployment bug affecting structured-output requests on GDPval-AA and AA-Briefcase. Anthropic states that the bug has since been fixed.
Benchmark results
Anthropic’s comparison includes Sonnet 5.5, Sonnet 5, Opus 5.5, and GPT-6 Sol, although not every model appears in every test.
On Terminal-Bench 4.0, Sonnet 5.5 scored 70.6%, compared with 10.3% for Sonnet 5 and 66.4% for Opus 5.5. On CursorBench 4.0, Sonnet 5.5 reached 55.5%, ahead of Sonnet 5 at 34.1% but behind Opus 5.5 at 57.8%.
The model also posted higher scores than Sonnet 5 on several knowledge-work and computer-use evaluations:
- GDPval-AA v2.1: 1,844 for Sonnet 5.5, compared with 1,449 for Sonnet 5 and 1,846 for Opus 5.5.
- AA-Briefcase v1.1: 1,811 for Sonnet 5.5, compared with 1,359 for Sonnet 5 and 1,822 for Opus 5.5.
- Humanity’s Last Exam with tools: 64.5% for Sonnet 5.5 versus 54.9% for Sonnet 5 and 67.7% for Opus 5.5.
- OSWorld 2.1 (partial): 80.1% for Sonnet 5.5 versus 57.0% for Sonnet 5 and 81.8% for Opus 5.5.
- Chartography without tools: 61.6% for Sonnet 5.5 versus 15.6% for Sonnet 5 and 64.4% for Opus 5.5.
On FrontierCode 1.1 (Main), Sonnet 5.5 scored 52.1% at Xhigh effort and 46.2% at Max effort. Sonnet 5 scored 42.4%, while Opus 5.5 reached 54.4% and GPT-6 Sol scored 49.3%.
Anthropic attributes the lower Max-effort result to more frequent code-review behavior that can cause timeouts or edits outside the task’s scope. The CursorBench comparison also uses GPT-5.6 Sol because GPT-6 Sol performance is not publicly reported for that benchmark. Anthropic further notes that a recently fixed OpenAI bug affecting image understanding may have influenced GPT-6 Sol’s AA-Briefcase, GDPval-AA, and Chartography results.
Design and safety changes
Beyond coding and general knowledge work, Anthropic positions Sonnet 5.5 as a model for producing polished interfaces and presentation materials. The company claims that it can add visual polish to user interfaces and follow templates closely enough to produce slides requiring minimal editing.
Anthropic also claims that Sonnet 5.5 writes more clearly than its previous-generation models and is suited to rapid iteration and collaboration on less complex tasks. Those descriptions are company assessments rather than independently verified findings in the supplied material.
Sonnet 5.5 is the first Sonnet model to include cyber safeguards and fallbacks similar to those used by Anthropic’s most capable models. Anthropic states that routine software development is unaffected. Its automated behavioral audit found that the model matched or improved on Sonnet 5 across most of the company’s measures for alignment and honesty.
Claude states that Sonnet 5.5 is available everywhere as of September 28, while Haiku 5.5 is scheduled to join the family in the coming weeks.
Source: Claude on X



