Sonnet 5.5 vs Opus 5.5: which coding tasks need Opus?
Choose Sonnet 5.5 or Opus 5.5 for bug fixes, code review and complex changes. Compare task scope, effort, cache costs and the evidence behind the decision.
Start by evaluating Sonnet 5.5 on well-defined coding tasks and Opus 5.5 on work where judgment, ambiguity or repeated failure justifies the extra trial. That is a workload-selection approach, not proof that Sonnet always handles small tasks or that Opus always wins on large ones. Your repository tests and review standards remain the acceptance criteria.
Anthropic positions Sonnet as a faster, lower-cost complement to Opus, while describing Opus as suited to complex work requiring careful judgment. The two models’ billing and effort settings make the choice less simple than “Sonnet is half the price.” This guide uses documentation checked September 29, 2026; no original head-to-head model trial is claimed.
Rates do not settle the choice
| Official Claude API category | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input, per million tokens | $2 | $4 |
| Output, per million tokens | $10 | $20 |
| Cache reads, per million tokens | $0.20 | $0.20 |
| API default effort | high | medium |
| Context window | 1M tokens | 1M tokens |
Sources: Sonnet specification and Opus specification. These are vendor rates, not subscription quotas or an Ofox quote. The cache-read row is especially relevant to repeated repository context: a large cached prefix does not have the same two-to-one difference as uncached input and output.
Imagine two synthetic requests with the same 100,000 cache-read tokens and 2,000 billable output tokens, with all other billable categories excluded. Sonnet would cost $0.04 and Opus $0.06. The difference is not twofold because cache reading costs the same in this example. If one model takes additional turns, the comparison changes again.
Match the task to a testable outcome
| Task | Useful starting comparison | Evidence to retain |
|---|---|---|
| Bug with a deterministic reproduction | Sonnet at a modest effort setting, then Opus if needed | Failing test, patch, full regression run |
| Small feature with explicit requirements | Sonnet versus your current working baseline | Acceptance checklist, scope of edits |
| Ambiguous cross-module failure | Include Opus from the outset | Competing explanations, inspected files, verified cause |
| Repository review | Give both the same scope | Confirmed findings versus false positives |
| Risky migration | Compare planning and validation separately | Migration checklist, rollback path, integration tests |
These are trial designs. They are not measured pass-rate claims. A short patch can require difficult reasoning, and a long but mechanical change can be easy. File count or lines changed alone is a poor measure of task difficulty.
Read benchmark claims with their conditions
Artificial Analysis’s Sonnet launch report shows strong results on several tasks but much higher output consumption at max effort. It also notes a pre-release structured-output issue affecting the tested deployment and planned reruns. Those results can justify testing both models; they cannot establish your cheapest configuration.
Do not compare an Opus medium run to a Sonnet max run and call the result a pure model comparison. It is a configuration comparison. That can still be valuable, but publish both settings and the actual budgets. Equal effort names also do not guarantee equal computation.
Anthropic’s own release announcement describes different model strengths and evaluation conditions. Keep vendor claims separate from independent measurements and from your own observations. A disagreement between charts may reflect tasks or settings rather than a mysterious contradiction.
A simple escalation policy
Before a run, write a stopping condition: for example, one proposed fix plus a regression check. If the check fails, inspect the failure before rerunning. If the model misunderstood the task, clarify the input rather than blindly increasing effort. If it found the right area but cannot produce a valid fix, trying Opus becomes a useful controlled escalation.
Carry the issue description, relevant files and test results into the new run. Do not assume encrypted thinking blocks are portable across models. Sonnet 5.5 has model- and conversation-specific rules described in the migration guide. Retain visible evidence rather than relying on hidden reasoning continuity.
Record the cost of the initial attempt and the escalation together. Otherwise a routing system can appear cheaper by attributing the first failure to one model and counting only the final successful run for another. Also record review time: a patch that compiles but needs substantial cleanup has not finished the job.
What to do on a subscription
Do not translate the API table into an exact number of Claude Code prompts. Subscription allowances and usage limits depend on the account and service conditions. Check the model actually selected, especially after a client update or provider change, and keep billing mode separate from the model name.
The Claude Code setup guide explains version checks and explicit selection. The effort guide separates API defaults from Claude Code defaults. If you are considering another provider, the Sonnet versus Sol guide gives a controlled comparison method.
Frequently Asked Questions
- Is Sonnet always half the cost of Opus?
- No. Its listed uncached input/output rates are half, but cache rates, token consumption, retries and tools affect completed-task cost. The worked example above deliberately isolates those differences.
- Should code review always use Opus?
- No universal rule follows from the documentation. Compare confirmed findings and false positives on representative reviews, including the time a human spends checking each finding.
- Can I move the same conversation between models?
- Visible messages can be part of a supported workflow, but thinking blocks have compatibility rules. Check the migration documentation rather than assuming the entire hidden state transfers.


