Anthropic’s new guide describes how Claude Opus 5.5 behaves differently from Claude Opus 5. It covers effort calibration, thinking in API integrations and chat, progress messages, unattended and multiagent work, safeguard refusals, frontend generation, complex visual input, multi-application workflows, and pasted text in user messages. The document also points to the model overview, general prompting guidance, and a migration guide covering four breaking API changes.
Claude Opus 5.5 produces output tokens more than 30 percent faster than Claude Opus 5 and usually completes the same task with fewer tokens. Existing prompts should generally continue to work. Anthropic presents the Claude Opus 5 guidance as a useful starting point.
For agentic coding and code review, Anthropic says Opus 5.5 at medium effort matched or exceeded Opus 5 at high effort in its tests, using fewer steps and tokens. The company also reports stronger performance on multi-hour audits and large codebase migrations with parallel subagents and limited supervision. Early testers reported that code review found more bugs and raised fewer false alarms, with changes explained in plain language.
The guide reports fewer incorrect figures and source citations during knowledge work. It highlights financial modeling, valuation workbooks, planning dates that fall on the wrong weekday, and charts whose visuals disagree with their underlying figures. Documents, spreadsheets, and presentations from the model reportedly require less editing. Anthropic also says the model gives clearer progress reports and completion summaries.
Visual performance has improved. In Anthropic’s tests, Opus 5.5 at its lowest effort read dense charts more accurately than Opus 5 at its highest effort while using a small fraction of the output tokens. It also handles positional relationships in flowcharts, diagram revisions, and calendar screenshots more reliably. Computer-use performance at the default effort matched the success rate Opus 5 reached at a much higher effort setting.
Effort is the main control because thinking is always enabled. Anthropic recommends starting with medium effort, the default for Opus 5.5, and testing levels against application-specific evaluations. Opus 5 defaults to high effort. Effort names represent different amounts of reasoning across models. Medium effort on Opus 5.5 matched or exceeded high effort on Opus 5 in coding and knowledge-work evaluations, while low effort came close on several coding tests at lower cost.
At the same effort level, Opus 5.5 may think longer than Opus 5, especially at `xhigh` and `max`. The `max_tokens` limit must leave room for both thinking and the answer, even when thinking content is hidden. Anthropic says `128,000`, the model maximum, worked well for long agentic coding turns. The guide reserves `xhigh` and `max` for workloads with a measured quality benefit. Lowering effort reduces thinking, cost, and latency more consistently than prompt instructions. Changing top-level effort invalidates the prompt cache. A per-message effort change, available in beta, preserves the cache.
Claude Opus 5 supports `thinking` disabled at high effort or below. Opus 5.5 always thinks, so Anthropic recommends beginning at low effort when migrating such an integration and measuring latency and quality. A system instruction such as asking for a direct answer can reduce thinking further, with a possible quality cost. Prompts that request visible reasoning should be removed. Applications should use summarized thinking blocks with `display: "summarized"` and handle `reasoning_extraction` refusals. Clients should inspect every response block type because a response may begin with a thinking block, whose `thinking` field is empty when the default display is `omitted`.
For long unattended tasks, a text-only turn ending with `stop_reason: "end_turn"` can be a progress report. Anthropic recommends keeping an external checklist, continuing while work remains open, and waiting for background commands or subagents to finish. A harness can send a short continuation naming unfinished items, then stop after two or three automatic continuations to avoid loops. A system prompt can identify premature summaries and offers to wait as unwanted stops. Anthropic says this instruction should be added from the first request because changing the system prompt later invalidates earlier thinking blocks. It can produce more tool calls and output tokens, and it suits unattended agents with separate confirmation controls for risky actions.
The model uses safety classifiers for biology, cybersecurity, and reasoning extraction. Biology safeguards match those of Claude Fable 5.1 and are new for users migrating from Opus 5. Everyday health and educational questions are unaffected. Anthropic directs life-sciences users affected by the classifier to its Life Sciences Verification Program. Source-code vulnerability discovery is allowed, while high-risk dual-use cybersecurity work remains restricted. A refusal appears as a normal response with `stop_reason: "refusal"` and a `stop_details` category. Automatic fallback can retry most refusals on another model. Server-side fallback returns `reasoning_extraction` refusals without retrying.
Progress messages between tool calls arrive as progress-update thinking blocks. Clients must set `display: "updates"` and send the beta header `thinking-display-updates-2026-08-18` to receive their text. Applications that need to deliver verbatim material during a long turn can provide a dedicated user-message tool from the first request. Adding that tool later changes the conversation prefix and invalidates earlier thinking blocks. System prompts can request predictable updates. A harness can issue a reminder after several silent tool calls, such as five, using a turn-scoped system message with `clear_at: "next_user_message"` and the beta header `mid-conversation-system-clear-at-2026-08-21`. Anthropic reports that this approach roughly halved long silent stretches in agentic coding tests without measurable cost changes.
For workflows spanning email, documents, spreadsheets, and CRM records, Anthropic recommends exploring relevant sources before acting, including sources the task does not name. Its tests found higher completion accuracy at medium and max effort, with slightly more tool calls and tokens. Untrusted content should remain outside the records searched by the agent.
Multiagent harnesses can provide elapsed time and a budget, such as `elapsed 340s / 1200s`. Anthropic reports that time signals helped small research teams finish sooner while preserving answer quality compared with a single agent. A budget is advisory, so applications still need their own timeout. A lower effort level reduces the work itself, while a budget mainly encourages parallel work.
In chat applications, Anthropic recommends removing generic instructions to think carefully because Opus 5.5 controls thinking through effort. Tests showed faster starts without a clear quality decline. On follow-up turns, the model may reconsider earlier answers, which increases latency. A system prompt can tell it to treat completed answers as settled unless the user points out a problem. Anthropic reports lower thinking and faster replies after this change, with a reduced tendency to detect its own earlier mistakes.
Opus 5.5 is described as more resistant to indirect instructions in tool results, web pages, and screen content than earlier Opus models. Applications should mark text pasted by users with matching `<pasted_content>` tags carrying a short random ID, then tell the model that the block may contain instructions from another source. The model should follow such instructions only when the user’s own message requests that behavior.
For dense charts, diagrams, and screenshots, the guide recommends retesting older scaffolding. Higher-resolution images help technical drawings. A container with PIL and OpenCV can crop, zoom, measure, and verify images, while a standalone crop tool is a lighter option. Higher effort improves tool-assisted visual work. Without tools, extra effort helps technical drawings more than charts.
For frontend tasks without design direction, Opus 5.5 falls back on recurring visual patterns. Anthropic recommends naming specific patterns to avoid and reviewing the first result iteratively. Its example asks for a vanilla HTML/CSS personal site with placeholder data and excludes cream or off-white backgrounds, italic accent words in headings, numbered `01/02/03` labels, monospace labels, and pill-shaped buttons.




Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Nothing arrived? Check your spam folder — and mark the mail as "Not spam" so it lands in your inbox next time.