TechCompare LogoTechCompare

Grok 4.6 API Pricing & Cost Calculator

Grok 4.6 at $2/$6 is priced like Qwen 3.8 Max, but xAI's 200K length tier is the trap: one long call bills entirely at $4/$12. Keep prompts under 200K and it's competitive with the Chinese flagships.

Grok 4.6 is xAI's current flagship API, keeping Grok 4.5's $2/$6 token rate but introducing a prompt-length tier: everything doubles once a request crosses 200K tokens, including cache reads.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.

How this is calculated

Grok 4.6 is priced at $2.00 per million input tokens and $6.00 per million output tokens for prompts under 200K tokens, and $4.00/$12.00 for prompts at or above 200K. Cache reads are $0.50/M under 200K and $1.00/M above. There is no published batch discount, and the full context window is 500K.

Verdict

A 100K + 5K call costs $0.23, matching Qwen 3.8 Max exactly at the low tier. Cross the 200K prompt threshold and the whole request bills at $4/$12, so a 250K + 5K call runs $1.06 versus $0.53 on a tierless $2/$6 model. Cache reads at $0.50/M (under 200K) are weaker than Alibaba's $0.20 and DeepSeek's $0.044, so stable-prompt loops cost more on xAI than any of those. The 500K window is the real differentiator if your workload sits between 200K and 1M.

More API Standalones scenarios

GPT-5.5 Pricing
Single-model gpt-5.5 cost estimate
View details ➜
GPT-5.4 Pricing
Single-model gpt-5.4 cost estimate
View details ➜
Claude Opus 4.8 Pricing
Single-model claude-opus-4.8 cost estimate
View details ➜

Frequently asked questions

How does the 200K threshold billing work on Grok 4.6?
If a request's prompt is below 200K tokens it bills at $2/M input and $6/M output. At or above 200K the entire request bills at $4/M input and $12/M output, including cache reads which go from $0.50/M to $1.00/M. There's no partial split - one long call pays the higher rate for all of it.
How does Grok 4.6 compare to Qwen 3.8 Max?
Same $2/$6 headline rate, but Qwen's cache reads are $0.20/M versus Grok's $0.50/M and Qwen has a 50% batch tier Grok lacks. Grok's edge is the 500K context (Qwen is 1M but with cheaper unit economics overall) and xAI's real-time integrations. On pure token cost, Qwen wins under 200K prompts unless you need what xAI uniquely offers.
How much does Grok 4.6 cost per million output tokens?
$6.00 per million below 200K-token prompts, $12.00 per million at or above. There's no batch discount. Cache reads are $0.50 per million in the low tier and $1.00 in the high tier.