Anthropic released Claude Opus 5.5 on September 22, 2026. It is the first model in the Claude 5.5 family. Anthropic places its performance close to Claude Fable 5.1 on most work and says typical operating costs are 40% below Claude Opus 5. The release follows CEO Dario Amodei’s call for pacing frontier AI development. Frontier Design and METR evaluated the model before launch.
Anthropic reports large gains on complex tasks. One tester completed a 680,000-line code migration in less than a day, work that the company says would have taken an engineering team weeks. Opus 5.5 reduced page load times in 39 of 40 web application tests. Opus 5 made smaller changes and also changed application behavior in those tests. A separate game-building test gave Opus 5.5 the strongest graphics and polish among the evaluated models.
The model uses less compute and generates output more than 30% faster than Opus 5. Anthropic’s typical-workload estimate puts the total cost reduction at 40%. Prices per million tokens are $4 for input, $20 for output, $0.20 for cache reads and $5 for cache writes. Opus 5 costs $5, $25, $0.50 and $6.25 for the same categories. Fast mode is available in Claude Code and the Claude Platform at up to 2.5 times the speed, for $8 per million input tokens and $40 per million output tokens. Anthropic is also raising five-hour limits for Pro, Max and Team plans and adding a rate-limit reset that subscribers can save for later use.
Coding is a primary focus. An early tester audited and fixed a 200,000-line codebase in under three hours. Opus 5 needed more than 20 hours and used 2.5 times as many tokens. In an internal translation of HAProxy from C to Rust, both Opus 5.5 and Fable 5.1 passed nearly all regression tests. Opus 5.5 completed the job in 9.5 hours, compared with 12 hours for Fable 5.1, and cost 51% less.
On Anthropic’s published benchmark table, Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1, 57.8% on CursorBench 4.0, 1846 Elo on GDPval-AA v2.1, 40.0% on AutomationBench, 67.7% with tools on Humanity’s Last Exam, 58.7% on Terminal-Bench-Science 0.1, 81.8% partial on OSWorld 2.0 and 89.0% with tools on Chartography. The corresponding Opus 5 scores were 52.3%, 48.0%, 46.6%, 1708 Elo, 26.9%, 63.6%, 29.0%, 74.0% and 83.4% where Anthropic reported a comparison.
At default effort, Opus 5.5 reached 54.6% on FrontierCode, ahead of GPT-6 Astra’s 53.3% top score at about one fifth of the cost per task. On CursorBench it reached 52.5%, compared with 51.8% for Fable 5.1 at maximum effort, 46.6% for Opus 5 and 41.7% for GPT-5.6 Sol. Anthropic says it beats GPT-5.6 Sol by 11 points at roughly one third of the cost per task. On Terminal-Bench 4.0 it matches GPT-6 Astra at about 40% of the cost.
The benchmark figures use adaptive thinking at maximum effort unless stated otherwise. Terminal-Bench uses xhigh effort for Opus 5.5 and high effort for GPT-6 Astra. Anthropic ran Opus 5.5 with production safeguards enabled. Safeguard interventions moved some cybersecurity tasks to Opus 4.8, biology tasks to Claude Opus 5 and frontier LLM development tasks to Claude Opus 5. Zapier’s AutomationBench run counted safeguard interventions as failures. Anthropic says these conditions may reduce the reported scores.
Early users reported fewer steps and tokens. GitHub saw Opus 5.5 solve more terminal tasks in VS Code with less than half the steps used by Opus 5. Clio ran a six-repository engineering task for more than 18 hours with minimal rework. Lovable reported one third to one half fewer steps. Quantium reduced a complex task from 38 prompts over four days to 11 prompts over three hours. Spotify, Optiver, Column and Kiro also reported lower token use or costs. Optiver measured 40% to 50% lower costs on comparable agentic coding work, while Kiro plans to add the model soon.
Anthropic also reports stronger knowledge-work performance. In a source-checking test, 16 of 18 Opus 5.5 reports met the quality bar, while Fable 5.1 and Opus 5 failed every attempt. Walleye Capital said the model found an off-by-one error in its evaluation instructions. In a fictional merger analysis, Opus 5.5 produced a more complete financial model and clearer presentation in 63 minutes, compared with 93 minutes for Opus 5, at half the production cost. On GDPval-AA v2.1, which covers 44 occupations, Opus 5.5 scored 1846 Elo, ahead of Fable 5.1 at 1735 and Opus 5 at 1708. It also led the reported AutomationBench and WANDR comparisons.
Customer evaluations supplied further figures. Deloitte reported that low-effort Opus 5.5 found 72% of known code-review bugs, compared with 56% for Opus 5 at high effort. Rogo said it beat Opus 5 on BigFinance Bench with about 60% fewer output tokens. Hex reported better detection of problems in delivery data. Hebbia measured 86.6% coverage on expert-graded finance workflows, compared with 60.3% for Opus 5. Viktor reported nearly half the cost and twice as many correctly solved hard tasks at the same effort.
Anthropic says the model’s writing is clearer and puts important information earlier. It follows supplied writing rules more reliably and uses less jargon. This is intended to improve long autonomous sessions and make outputs easier to review. Claude Sonnet 5.5 and Claude Haiku 5.5 are planned for the following weeks.
Opus 5.5 includes a classifier that checks every action before execution, an open-source sandbox for security review and code review intended to catch vulnerabilities before merging. Anthropic says the model matched or exceeded Opus 5 against prompt injection in coding, tool use, computer use and web browsing. Gray Swan’s evaluation gave Opus 5.5 and Fable 5.1 the lowest prompt-injection success rate among the tested models.
Anthropic’s automated behavioral audit covers nearly 2,000 simulated scenarios. Opus 5.5 achieved the company’s strongest results on nearly every recent measure of misaligned behavior and on most honesty measures. In a containment-boundary test, it attempted circumvention about 85% less often than Opus 5 or Claude Mythos 5.1. Every attempt was classified as low severity and self-reported. Anthropic says evaluations remain incomplete because the model can often recognize that it is being tested.
The company applies Fable 5.1-class safeguards to Opus 5.5 for cybersecurity, biology and distillation. Most cybersecurity tasks are routed transparently to Opus 4.8. Anthropic plans to extend its Cyber Verification Program with three access tiers, including access to Claude Mythos models. Opus 5.5 matched or exceeded Claude Mythos 5.1 in several biology evaluations, including a long-horizon molecular prediction and design test conducted with Dyno Therapeutics. Vetted organizations can apply to the Life Sciences Verification Program.
Anthropic’s September 2026 threat intelligence report describes detected distillation activity. Opus 5.5 launches with preserved thinking, which prevents API users from editing Claude’s earlier context to extract reasoning. The restriction applies to Fable 5.1 and Opus 5.5 API accounts created on or after August 31, 2026. The model supports zero data retention and includes watermarking measures for the EU AI Act. Thinking mode cannot be disabled.
Anthropic says its current safety work combines alignment testing, outside evaluation by METR and Frontier Design, capability-based safeguards and reporting under its Responsible Scaling Policy. Its future work includes stricter reinforcement-learning environment filters, improved alignment rewards, automatically generated safety scenarios, stronger security monitoring and interpretability-based evaluation. The company says models capable of automating AI research will require higher standards and broader public-policy involvement. It also points to its work with Accenture on embedded evaluation.
Claude Opus 5.5 is available through Claude, Claude Code and the Claude Platform. Developers can use the model identifier claude-opus-5-5. The model is also available through Amazon Web Services, Google Cloud and Microsoft Azure.




Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Nothing arrived? Check your spam folder — and mark the mail as "Not spam" so it lands in your inbox next time.