TechCompare LogoTechCompare

DeepSeek V4 Pro API Pricing & Cost Calculator

DeepSeek V4 Pro at $1.32/$3.96 peak is the flagship-tier reasoning API priced like a workhorse. It's an order of magnitude cheaper than GPT-5.6 Sol ($4/$20) on output, with a 96.7% cache discount on top.

DeepSeek V4 Pro's 0813 refresh is the GA version of the trillion-parameter reasoning flagship, now 17.5% cheaper than the spring preview at $1.32/$3.96 peak with the same aggressive cache discount.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

How this is calculated

DeepSeek V4 Pro is priced at $1.32 per million input tokens and $3.96 per million output tokens at peak, with off-peak rates of $0.66/$1.98 - half price outside the 01:00-04:00 and 06:00-10:00 UTC windows. Cached input at peak is $0.044 per million, a 96.7% discount. Context is 1M tokens with a 384K max output.

Verdict

The per-call math at 100K + 5K is $0.15 peak, $0.08 off-peak - versus $0.50 for GPT-5.6 Sol and $0.23 for Qwen 3.8 Max. That's over 3x cheaper than OpenAI's flagship with comparable benchmark ceilings on most reasoning work. The 96.7% cache discount drops cached input to $0.044/M, so agent loops with a stable system prompt see input cost collapse. The April preview was $1.60/$3.20, so August's $1.32/$3.96 card is a quiet 17.5% input cut that rewards keeping the integration current.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
MiMo-V2.6-Flash Pricing
Cost calculator for this model
View details ➜
GLM-5.3-Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

Is the 0813 refresh priced differently from the spring V4 Pro preview?
Yes. The April preview was $1.60/M input and $3.20/M output. The August 0813 GA refresh is $1.32/M input and $3.96/M output, so input dropped 17.5% while output rose slightly. Cache-hit pricing stays at the same deep discount tier.
DeepSeek V4 Pro or Qwen 3.8 Max?
V4 Pro is $1.32/$3.96 peak versus Qwen 3.8 Max's $2.00/$6.00, and DeepSeek's cache discount is deeper (96.7% vs 90%). Qwen's 50% batch tier partially closes the gap for offline jobs. For the hardest reasoning benchmarks the two trade wins depending on the eval - pick V4 Pro on raw price and caching, Qwen if you're in the Alibaba ecosystem.
What does off-peak scheduling save?
Exactly half. V4 Pro off-peak is $0.66/M input and $1.98/M output versus $1.32/$3.96 peak. Peak windows are 01:00-04:00 and 06:00-10:00 UTC, so anything that queues outside those hours doubles your token budget for free.