DeepSeek V4 Pro API Pricing & Cost Calculator
DeepSeek V4 Pro at $1.32/$3.96 peak is the flagship-tier reasoning API priced like a workhorse. It's an order of magnitude cheaper than GPT-5.6 Sol ($5/$30) on output, with a 96.7% cache discount on top.
DeepSeek V4 Pro's 0813 refresh is the GA version of the trillion-parameter reasoning flagship, now 17.5% cheaper than the spring preview at $1.32/$3.96 peak with the same aggressive cache discount.
By TechCompare · Updated
Calculator
Cost Comparison
Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.
How this is calculated
DeepSeek V4 Pro is priced at $1.32 per million input tokens and $3.96 per million output tokens at peak, with off-peak rates of $0.66/$1.98 - half price outside the 01:00-04:00 and 06:00-10:00 UTC windows. Cached input at peak is $0.044 per million, a 96.7% discount. Context is 1M tokens with a 384K max output.
Verdict
The per-call math at 100K + 5K is $0.15 peak, $0.08 off-peak - versus $0.65 for GPT-5.6 Sol and $0.23 for Qwen 3.8 Max. That's 4x cheaper than OpenAI's flagship with comparable benchmark ceilings on most reasoning work. The 96.7% cache discount drops cached input to $0.044/M, so agent loops with a stable system prompt see input cost collapse. The April preview was $1.60/$3.20, so August's $1.32/$3.96 card is a quiet 17.5% input cut that rewards keeping the integration current.
More API Standalones scenarios
Related guides
Frequently asked questions
Is the 0813 refresh priced differently from the spring V4 Pro preview?
DeepSeek V4 Pro or Qwen 3.8 Max?
What does off-peak scheduling save?
Related tools
LLM VRAM Calculator
Calculate the VRAM needed to run or fine-tune any LLM at any quantization.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜JSON Formatter
Validate, format, and minify JSON data with syntax highlighting.
Use tool ➜