TechCompare LogoTechCompare

DeepSeek V4 Flash API Pricing & Cost Calculator

DeepSeek V4 Flash is the cheapest 1M-context frontier API you can buy. Cache-hit input at $0.014/M makes stable-prompt loops nearly free, and scheduling off-peak halves the entire bill.

DeepSeek V4 Flash is the small fast member of the V4 pair, priced for high-volume traffic with a peak/off-peak rate split and one of the deepest prompt-cache discounts of any frontier vendor.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
1,000 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 1,000 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

How this is calculated

DeepSeek V4 Flash is priced at $0.44 per million input tokens and $1.32 per million output tokens at peak hours (01:00-04:00 and 06:00-10:00 UTC), dropping to $0.22/$0.66 off-peak. Cached input at peak is $0.014 per million, a 96.8% discount. The context window is 1M tokens with up to 384K output.

Verdict

Peak rates land a 100K + 5K call at $0.05, which no US lab comes close to - GPT-5.6 Sol runs $0.50 for the same call. The real trick is the two levers stacked: cache-hit input at $0.014/M drops the input side by 30x on repeated context, and off-peak scheduling halves both rows for any workload that can wait. Against DeepSeek V4 Pro 0813 ($1.32/$3.96) Flash is 3x cheaper per token and is the right default unless the job genuinely needs Pro's deeper reasoning.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
MiMo-V2.6-Flash Pricing
Cost calculator for this model
View details ➜
GLM-5.3-Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

What do DeepSeek's peak and off-peak hours mean for billing?
Peak is 01:00-04:00 and 06:00-10:00 UTC. V4 Flash at peak is $0.44/M input and $1.32/M output. Off-peak halves both to $0.22/M input and $0.66/M output. Batch-style workloads that can queue overnight effectively cost half the published rate.
How deep is the DeepSeek prompt cache discount?
96.8% on V4 Flash. Cache-hit input is $0.014 per million versus $0.44 standard. For agent loops that reuse the same system prompt and tool definitions, the input side of the bill nearly disappears. V4 Pro has a similar 96.7% discount ($1.32 to $0.044).
DeepSeek V4 Flash or V4 Pro?
Flash for volume: it's 3x cheaper per token ($0.44/$1.32 vs $1.32/$3.96 peak) with the same 1M context. Pro for hard reasoning, multi-step planning, and anything where the evals clearly separate. Mix them by routing - Flash as default, Pro as escalation.