TechCompare LogoTechCompare

DeepSeek V4 Pro API Pricing & Cost Calculator

DeepSeek V4 Pro at $1.32/$3.96 peak is the flagship-tier reasoning API priced like a workhorse. It's an order of magnitude cheaper than GPT-5.6 Sol ($5/$30) on output, with a 96.7% cache discount on top.

DeepSeek V4 Pro's 0813 refresh is the GA version of the trillion-parameter reasoning flagship, now 17.5% cheaper than the spring preview at $1.32/$3.96 peak with the same aggressive cache discount.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.

How this is calculated

DeepSeek V4 Pro is priced at $1.32 per million input tokens and $3.96 per million output tokens at peak, with off-peak rates of $0.66/$1.98 - half price outside the 01:00-04:00 and 06:00-10:00 UTC windows. Cached input at peak is $0.044 per million, a 96.7% discount. Context is 1M tokens with a 384K max output.

Verdict

The per-call math at 100K + 5K is $0.15 peak, $0.08 off-peak - versus $0.65 for GPT-5.6 Sol and $0.23 for Qwen 3.8 Max. That's 4x cheaper than OpenAI's flagship with comparable benchmark ceilings on most reasoning work. The 96.7% cache discount drops cached input to $0.044/M, so agent loops with a stable system prompt see input cost collapse. The April preview was $1.60/$3.20, so August's $1.32/$3.96 card is a quiet 17.5% input cut that rewards keeping the integration current.

More API Standalones scenarios

GPT-5.5 Pricing
Single-model gpt-5.5 cost estimate
View details ➜
GPT-5.4 Pricing
Single-model gpt-5.4 cost estimate
View details ➜
Claude Opus 4.8 Pricing
Single-model claude-opus-4.8 cost estimate
View details ➜

Frequently asked questions

Is the 0813 refresh priced differently from the spring V4 Pro preview?
Yes. The April preview was $1.60/M input and $3.20/M output. The August 0813 GA refresh is $1.32/M input and $3.96/M output, so input dropped 17.5% while output rose slightly. Cache-hit pricing stays at the same deep discount tier.
DeepSeek V4 Pro or Qwen 3.8 Max?
V4 Pro is $1.32/$3.96 peak versus Qwen 3.8 Max's $2.00/$6.00, and DeepSeek's cache discount is deeper (96.7% vs 90%). Qwen's 50% batch tier partially closes the gap for offline jobs. For the hardest reasoning benchmarks the two trade wins depending on the eval - pick V4 Pro on raw price and caching, Qwen if you're in the Alibaba ecosystem.
What does off-peak scheduling save?
Exactly half. V4 Pro off-peak is $0.66/M input and $1.98/M output versus $1.32/$3.96 peak. Peak windows are 01:00-04:00 and 06:00-10:00 UTC, so anything that queues outside those hours doubles your token budget for free.