TechCompare LogoTechCompare

Claude Sonnet 5.5 API pricing and cost calculator

A Sonnet request with 100K input and 5K output costs $0.25 uncached or $0.16 with half the input cached. Prompt growth raises token volume without adding a long-context price tier.

Claude Sonnet 5.5 charges $2 per million input tokens and $10 per million output tokens at standard rates. Its $0.20 cache-read price reduces repeated-context spending, while standard token rates stay flat within its 1M context window. Those properties are useful when budgeting document processing or ongoing conversations. The bill still depends on how much new input and output each turn adds, so start with actual usage rather than a model's context capacity.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

How this is calculated

At 100,000 input and 5,000 output tokens, Sonnet costs $0.20 for uncached input plus $0.05 for output, totaling $0.25. The starting calculator example reads half the input from cache, giving $0.10 uncached input, $0.01 cached input, and $0.05 output. That is $0.16 per request or $16 for 100 monthly requests. A 300,000-input, 5,000-output prompt costs $0.65 with no cache hits or $0.38 with half the input cached. A longer prompt doesn't introduce a separate token-rate tier within the advertised window. These examples exclude cache writes and tool charges. The embedded estimate uses OpenRouter's live provider rates, so check its routing assumptions alongside Anthropic's standard price list.

Verdict

Sonnet is worth testing when its results meet your quality targets and a predictable token rate helps you budget long prompts. Keep output concise when the task allows it, and compare retrieval with sending the full document on every turn. Prompt caching is valuable for reusable prefixes, but count the writes and misses too. If another model completes a task in fewer calls, its total session cost can matter more than a small difference in input rates.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
MiMo-V2.6-Flash Pricing
Cost calculator for this model
View details ➜
GLM-5.3-Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

What is the cost of 10,000 Sonnet 5.5 requests?
For the 100,000-input, 5,000-output example, 10,000 requests cost $2,500 with no cache hits or $1,600 with half the input cached. Those are token estimates, excluding writes, tool charges, and retries. For a real budget, group requests by their actual prompt lengths and cache usage instead of assuming every request looks like the same example.
How much does Sonnet 5.5 prompt caching cost?
Reads cost $0.20 per million tokens, a 90% reduction from the uncached input rate. Five-minute writes cost $2.50 per million tokens and one-hour writes cost $4. Cache reuse needs to recover those creation costs. Estimate the number of hits before expiration and avoid treating an input prefix that changes each turn as a guaranteed read hit.
Is Sonnet 5.5 cheaper than GPT-6.1 Sol?
Their base input and output rates match. Sol has a lower cache-read price at its base tier, while Sonnet avoids Sol's prompt-length surcharge within its own context window. Which costs less therefore depends on prompt length, cache usage, and output volume. The two providers can also bill different token counts for the same text, so measure usage on your own requests.
Does Sonnet's 1M context window mean I should use it all?
No. Capacity is a limit, not a spending target. Sending more material raises input volume even when the per-token rate stays flat. Keep relevant document sections and tool results, remove needless repetition, and test whether the shorter prompt preserves task quality. A targeted prompt may cost less and be easier to evaluate than a large collection sent on every turn.
What does Sonnet 5.5 cost in batch mode?
Anthropic's Batch API halves applicable token prices. The uncached example drops from $0.25 to $0.125 per request, and the half-cached token example drops from $0.16 to $0.08 before write costs and tools. That is an asynchronous provider workflow. The calculator models those rates when batch mode is enabled, rather than claiming OpenRouter's live requests have the same discount.