TechCompare LogoTechCompare

Grok 4.6 API Pricing & Cost Calculator

Grok 4.6 at $2/$6 is priced like Qwen 3.8 Max, but xAI's 200K length tier is the trap: one long call bills entirely at $4/$12. Keep prompts under 200K and it's competitive with the Chinese flagships.

Grok 4.6 is xAI's current flagship API, keeping Grok 4.5's $2/$6 token rate but introducing a prompt-length tier: everything doubles once a request crosses 200K tokens, including cache reads.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

How this is calculated

Grok 4.6 is priced at $2.00 per million input tokens and $6.00 per million output tokens for prompts under 200K tokens, and $4.00/$12.00 for prompts at or above 200K. Cache reads are $0.50/M under 200K and $1.00/M above. There is no published batch discount, and the full context window is 500K.

Verdict

A 100K + 5K call costs $0.23, matching Qwen 3.8 Max exactly at the low tier. Cross the 200K prompt threshold and the whole request bills at $4/$12, so a 250K + 5K call runs $1.06 versus $0.53 on a tierless $2/$6 model. Cache reads at $0.50/M (under 200K) are weaker than Alibaba's $0.20 and DeepSeek's $0.044, so stable-prompt loops cost more on xAI than any of those. The 500K window is the real differentiator if your workload sits between 200K and 1M.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
MiMo-V2.6-Flash Pricing
Cost calculator for this model
View details ➜
GLM-5.3-Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

How does the 200K threshold billing work on Grok 4.6?
If a request's prompt is below 200K tokens it bills at $2/M input and $6/M output. At or above 200K the entire request bills at $4/M input and $12/M output, including cache reads which go from $0.50/M to $1.00/M. There's no partial split - one long call pays the higher rate for all of it.
How does Grok 4.6 compare to Qwen 3.8 Max?
Same $2/$6 headline rate, but Qwen's cache reads are $0.20/M versus Grok's $0.50/M and Qwen has a 50% batch tier Grok lacks. Grok's edge is the 500K context (Qwen is 1M but with cheaper unit economics overall) and xAI's real-time integrations. On pure token cost, Qwen wins under 200K prompts unless you need what xAI uniquely offers.
How much does Grok 4.6 cost per million output tokens?
$6.00 per million below 200K-token prompts, $12.00 per million at or above. There's no batch discount. Cache reads are $0.50 per million in the low tier and $1.00 in the high tier.