SikiT

Field notes on running AI agents without a human in the loop.

Field notes on unattended AI agents.

7–11 minutes

Claude Opus 5.5 vs. Sonnet 5.5: which model fits your work?

Compare Claude Opus 5.5 and Sonnet 5.5 using official benchmarks and API prices, then choose a model and effort level for coding, documents, research, and agents.

Sources are linked in this article. Found an error? Report a correction.

Use Claude Sonnet 5.5 for well-defined coding, document, and automation tasks; choose Claude Opus 5.5 when success depends on resolving ambiguity, weighing conflicting evidence, or sustaining judgment across a long project. This guide uses official information checked on October 3, 2026, and separates vendor-reported benchmarks from practical recommendations; it does not claim a head-to-head test on your workload.

On this page

Recommendation: start with Sonnet for scoped work

Sonnet 5.5 at medium effort is a useful starting point for a task with a clear deliverable and a checkable result. Use Opus 5.5 at medium effort when the difficult part is deciding what to do: tracing a failure across services, selecting an architecture under conflicting constraints, or reconciling sources that disagree.

That recommendation follows from the reported performance, pricing, and effort guidance below. It is an editorial starting point, not a measured guarantee that Sonnet will be sufficient. Anthropic’s general model overview recommends starting with Opus 5.5 for most workloads; this guide starts with Sonnet for specifically bounded work where cost and response time matter.

Opus 5.5 was released on September 22, 2026; Sonnet 5.5 followed on September 28, 2026. Both are current models, so the decision is about workload fit rather than whether to adopt an older generation.

Decision criteria: difficulty, effort, and cost per finished task

Judge difficulty by the uncertainty the model must resolve. A large, repetitive extraction job may suit Sonnet even if the input is long. A short incident report with contradictory logs may justify Opus. Context capacity alone does not settle that choice: both models advertise a 1 million-token context window and a standard maximum output of 128,000 tokens.

Effort controls how much reasoning and output the model spends on a response. Anthropic recommends medium for Sonnet’s well-specified agentic coding and multistep tool use, high for harder or longer work, and medium or low for latency-sensitive chat. The Claude API defaults differ: Sonnet uses high, while Opus uses medium. Set effort explicitly when comparing them so an unnoticed default does not decide your experiment.

Track the cost of an accepted result: model usage, tool charges, retries, and the time required to check or repair the output. Sonnet’s lower token rates help only if its complete workflow meets your quality bar. Opus can be the economical choice when it avoids enough failed attempts or manual repair; that must be measured on your tasks.

Important differences: benchmark scores and API costs

What the published scores show

These are Anthropic-reported results from the Sonnet 5.5 system card and launch comparison, not SikiT measurements. Higher is better within each row; scores across different benchmarks cannot be combined into a meaningful overall percentage.

EvaluationOpus 5.5Sonnet 5.5What it evaluates
Terminal-Bench 4.066.4%70.6%Multistep terminal tasks; Opus result uses xhigh
FrontierCode 1.1, main set54.4%52.1% at xhigh; 46.2% at maxWhether code changes are mergeable
CursorBench 4.057.8%55.5%Ambiguous coding tasks across files
GDPval-AA v2.11846 Elo1844 EloProfessional knowledge work
AA-Briefcase v1.11822 Elo1811 EloLong-horizon knowledge work
Humanity’s Last Exam, with tools67.7%64.5%Multidisciplinary reasoning
OSWorld 2.1, partial setting81.8%80.1%Computer use in the reported setting

The system card’s default evaluation setting is max effort. Terminal-Bench uses Opus at xhigh, and the FrontierCode row includes Sonnet’s xhigh alternative. These scores do not describe both models at the medium settings recommended for starting work.

Sonnet leads this Terminal-Bench comparison and is close on several other rows. Anthropic nevertheless reports Opus as stronger on complex, open-ended work requiring sustained judgment. Treat that as a vendor assessment, not evidence of a fixed advantage on every job.

The effort settings are not identical throughout. Sonnet’s FrontierCode result also drops at max effort: the launch footnote describes extra review work causing timeouts or out-of-scope edits. A higher effort setting can therefore reduce task success.

What the API prices mean

Standard Claude API prices below are in US dollars per million tokens. They are usage rates, not Claude subscription prices or a forecast of how many chats your plan includes.

ItemOpus 5.5Sonnet 5.5
Uncached input$4$2
Output$20$10
Cache write, 5 minutes$5$2.50
Cache write, 1 hour$8$4
Cache read$0.20$0.20
Context window1 million tokens1 million tokens
Standard maximum output128,000 tokens128,000 tokens
Claude API default effortmediumhigh
Model IDclaude-opus-5-5claude-sonnet-5-5

For an illustrative request with 100,000 uncached input tokens and 10,000 billed output tokens, Sonnet costs $0.30 and Opus $0.60. This calculation excludes tools, retries, cache operations, and other pricing modifiers. It assumes identical billed token volumes, which real sessions may not use.

Workloads that reuse cached input need a separate calculation because the cache-read rates are equal. Anthropic also offers a 50% Batch API discount on input and output; provider-specific prices, fast mode, and data-residency charges can change the bill.

Anthropic’s claim that Sonnet 5.5 generates output over 30% faster compares it with Sonnet 5. It does not establish a 30% speed advantage over Opus 5.5 or over a complete workflow that waits on tools.

Choose by your situation: model and effort for each task

The choices below are practical recommendations inferred from the official evidence. The final column names the failure or decision that should prompt you to reconsider the starting choice.

Your taskStarting choiceReconsider when
Email drafts, summaries, translation, or editing supplied textSonnet, low or mediumThe result changes meaning or loses a material qualification; improve the brief and compare a reviewed Opus result
A bug fix with a reproduction and explicit acceptance checksSonnet, mediumThe cause crosses services or repeated attempts fail the same check; use Opus, medium or high
UI implementation or slides from an existing design and templateSonnet, mediumRequirements conflict or information hierarchy needs substantial redesign; ask Opus to resolve those decisions
Architecture, a migration plan, or an unfamiliar failureOpus, medium; high when neededThe design is settled and implementation can be split into bounded, testable tasks for Sonnet
Extracting facts from a supplied source packetSonnet, mediumSources disagree, citations do not support claims, or the conclusion requires weighing evidence; use Opus
A spreadsheet or business reportSonnet, mediumAssumptions interact across models or the recommendation requires sensitivity analysis; use Opus and independently recalculate
A long-running agent projectOpus for uncertain planning; Sonnet for scoped implementationA worker cannot show the requested result; return the failure evidence to the planner
Repetitive, high-volume automationSonnet, low or medium after evaluationQuality failures or retries erase the saving; compare accepted-result cost with Opus

Give either model a verifiable finish line

For a scoped coding task, provide the failing example, files it may change, expected behavior, and the check that proves completion. For a document task, provide the source material, intended reader, format, and facts that must survive the rewrite. Changing models cannot supply missing requirements or unavailable sources.

A useful implementation prompt is:

Fix the reported failure within the stated scope. Use the supplied reproduction and acceptance checks. Report what changed and the observed check result. Ask for clarification only when a missing requirement changes the implementation; obtain approval before any separately restricted action.

If Sonnet stops before finishing a multipart task at low or medium effort, Anthropic’s prompting guide recommends trying a higher effort level. If the task itself needs broader judgment, compare Opus at medium rather than repeatedly escalating Sonnet to max. Hand over the exact failure, source packet, and acceptance criteria so the next attempt can address the cause.

Apply the choice in Claude Code or the API

On the first-party Anthropic connection, Claude Code’s sonnet and opus aliases currently resolve to these 5.5 models. Other providers can resolve those aliases to older versions. Check the displayed model, or select the appropriate full model ID for your provider. Current documentation requires Claude Code v2.1.284 or later for Sonnet 5.5 and v2.1.280 or later for Opus 5.5.

Run /model to open the picker. For a temporary comparison, pressing s on a model’s row changes this session only; accepting with Enter saves the choice for future sessions. Claude Code also provides opusplan, which uses Opus in plan mode and Sonnet for execution, subject to the provider’s alias resolution.

Use /effort to select the effort level, and check the session header for the level actually in effect. Organization policies can cap the available setting.

For Claude API requests, use the exact model ID in the table and set output_config.effort explicitly. Opus 5.5 always uses adaptive thinking. Sonnet 5.5’s lowest thinking mode is between_tools, accepted at high effort or below; old integrations that send disabled thinking need migration checks.

Compare outcomes before committing a workflow

Use the same representative tasks, inputs, tool permissions, and acceptance checks for both models. Record the model and effort, actual elapsed time, billed usage, retries, and manual corrections. For code, run the reproduction and relevant tests; for documents, verify facts and inspect the rendered output. Prefer the less costly setting that reliably passes those checks, and keep a separate route for the difficult cases it fails.

Limits and evidence: what this comparison can establish

Official sources were checked on October 3, 2026. The benchmark table reports selected settings and workloads, and several results are close; it does not establish statistical superiority or a universal winner. Anthropic notes that the Sonnet deployment used for GDPval-AA and AA-Briefcase had a structured-output bug that was later fixed, so the published figures also describe a particular evaluation deployment.

This article has not measured the two models on your repository, writing style, language mix, or account. Subscription availability and usage limits depend on the account and provider. A long context window, fluent answer, or second model agreeing with the first does not verify completion. Before adopting the recommendation, confirm that the chosen model finishes your representative task and that its output passes the same acceptance checks.

Sources

Related articles

Stay in the loop

Get new agent guides with their test conditions and limitations. Unsubscribe anytime.

Comments

Questions, corrections, and useful counterpoints are welcome. Keep comments specific and on topic.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Thanks for commenting

Get new agent guides with their test conditions and limitations. Unsubscribe anytime.

Return to the comments