TechCompare LogoTechCompare

Claude Opus 5.5 vs Claude Sonnet 5.5: what the price difference buys

Sonnet is cheaper for the same uncached input and output. Cache reads tie, so high prompt reuse narrows the total gap without removing Opus's higher output cost.

Opus doubles the base input and output rates, while cache reads cost the same on both.

Claude Opus 5.5 charges twice Claude Sonnet 5.5's base input and output rates. The cache-read price is an exception: both charge $0.20 per million tokens read from cache. That makes a shared-context workflow different from a series of unrelated prompts. The right question is how much uncached input and output your application pays for, and whether using Opus improves the completed result enough to cover those higher rates.

By TechCompare · Updated

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

Option A
Claude Opus 5.5
Wins 0 of 10 compared specs
Option B
Claude Sonnet 5.5
Wins 7 of 10 compared specs

Side-by-side specs

SpecClaude Opus 5.5Claude Sonnet 5.5
Standard input per 1M tokens
$4.00
$2.00 (better on this spec)
Standard output per 1M tokens
$20.00
$10.00 (better on this spec)
Cache reads per 1M tokens
$0.20
$0.20
5-minute cache writes per 1M tokens
$5.00
$2.50 (better on this spec)
1-hour cache writes per 1M tokens
$8.00
$4.00 (better on this spec)
100K input + 5K output, no cache hits
$0.50
$0.25 (better on this spec)
100K input + 5K output, 50% cached input
$0.31
$0.16 (better on this spec)
100K input + 5K output, all input cached
$0.12
$0.07 (better on this spec)
Context window
1,000,000 tokens
1,000,000 tokens
Provider Batch API discount
50%
50%

How they differ

Opus costs $4 input and $20 output per million tokens, compared with Sonnet's $2 and $10. A 100,000-input, 5,000-output request without cache hits costs $0.50 on Opus and $0.25 on Sonnet. With half the input read from cache, the estimates become $0.31 and $0.16. If all 100,000 input tokens are eligible cache reads, they become $0.12 and $0.07. The gap narrows because cached input costs the same, but output remains twice as expensive on Opus. Five-minute cache writes cost $5 per million on Opus and $2.50 on Sonnet, with one-hour writes at $8 and $4. Both support a 1M context window at standard rates and a 50% provider Batch API discount. These worked estimates exclude writes and tool fees. The embedded chart uses live OpenRouter prices and can differ with provider routing.

Verdict

Use Sonnet as a candidate for tasks it completes to your required standard, then evaluate Opus on the work that remains difficult. An agent with frequent cache hits can make Opus less than twice as expensive per call, but that doesn't automatically make it better value. Track the number of attempts, output tokens, and human corrections per successful task. Route using those measurements rather than treating the Opus name as evidence of a benefit on every workload.

Which should you pick?

Choose Claude Opus 5.5

Use Opus when your evaluation shows a worthwhile improvement on longer agent workflows, coding tasks, or document analysis. A reused cache can reduce the cost of repeated context, but creating and refreshing that cache still costs money. Set budgets for reasoning and tool turns so a small quality gain doesn't arrive with an unexpectedly large output bill.

Choose Claude Sonnet 5.5

Use Sonnet when it reaches the same success criteria with acceptable response times and review effort. Its lower uncached-input, output, and write rates matter for one-off requests, growing conversations, and output-heavy applications. Keep an escalation path for tasks that need another attempt on Opus instead of paying the higher rate for every routine request.

Related comparisons

MiMo-V2.6-Pro vs MiMo-V2.6-Flash
Compare Xiaomi's normal-speed real-time rates, cache savings, and the cost of successful tasks.
Read comparison ➜
GLM-5.3-Flash vs DeepSeek V4.1 Flash
Cache hits can change the cheaper option off-peak. Compare Z.ai and DeepSeek's direct rates.
Read comparison ➜
MiMo-V2.6-Flash vs DeepSeek V4.1 Flash
Compare Xiaomi's flat real-time rates with DeepSeek's peak and off-peak token prices.
Read comparison ➜
GPT-6 Luna vs GPT-6.1 Sol
Luna's base input and output rates are 20 times lower. Compare the cost of a successful result.
Read comparison ➜
GPT-6 Luna vs Claude Haiku 4.5
Luna's base token rates are ten times lower. Task quality and integration decide whether switching pays off.
Read comparison ➜
GPT-6.1 Sol vs Claude Opus 5.5
Sol costs half as much at base rates, but its long-prompt surcharge narrows the gap.
Read comparison ➜

Frequently asked questions

Does Opus 5.5 always cost twice as much as Sonnet 5.5?
No. The twofold difference applies to uncached input, output, and cache writes, but cache reads are $0.20 per million tokens on both. A request composed largely of cached input has a smaller total price ratio. Opus can also produce a different number of output tokens or require a different number of calls, so actual task costs need measurement.
What does the same cache-read price mean for an agent?
An agent that reads an already-cached 100,000-token prefix pays $0.02 for that prefix on either model. It still pays for newly added input and generated output, where Sonnet's rates are lower. Don't assume that every tool result or message is a cache hit. Cache creation, expiration, and changing prefixes affect how much of a session qualifies for read pricing.
How much do 10,000 calls cost with half the input cached?
With 100,000 input and 5,000 output tokens per call, the token estimate is $3,100 on Opus and $1,600 on Sonnet. That's a $1,500 difference before cache writes and tools. The example assumes the same usage mix on every call. A production budget should account for short calls, large calls, retries, and changes in output length.
Should I pick Opus just because the task is important?
Importance alone doesn't establish which model completes a task more reliably. Compare both on representative cases and measure errors that matter to your application. Use Opus where the improvement is worth its additional cost. Keep validation and human review where needed on either model rather than assuming that paying a higher token rate guarantees a correct result.