DeepSeek V4 Flash API Pricing & Cost Calculator
DeepSeek V4 Flash is the cheapest 1M-context frontier API you can buy. Cache-hit input at $0.014/M makes stable-prompt loops nearly free, and scheduling off-peak halves the entire bill.
DeepSeek V4 Flash is the small fast member of the V4 pair, priced for high-volume traffic with a peak/off-peak rate split and one of the deepest prompt-cache discounts of any frontier vendor.
By TechCompare · Updated
Calculator
Cost Comparison
Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 1,000 requests.
How this is calculated
DeepSeek V4 Flash is priced at $0.44 per million input tokens and $1.32 per million output tokens at peak hours (01:00-04:00 and 06:00-10:00 UTC), dropping to $0.22/$0.66 off-peak. Cached input at peak is $0.014 per million, a 96.8% discount. The context window is 1M tokens with up to 384K output.
Verdict
Peak rates land a 100K + 5K call at $0.05, which no US lab comes close to - GPT-5.6 Sol runs $0.65 for the same call. The real trick is the two levers stacked: cache-hit input at $0.014/M drops the input side by 30x on repeated context, and off-peak scheduling halves both rows for any workload that can wait. Against DeepSeek V4 Pro 0813 ($1.32/$3.96) Flash is 3x cheaper per token and is the right default unless the job genuinely needs Pro's deeper reasoning.
More API Standalones scenarios
Related guides
Frequently asked questions
What do DeepSeek's peak and off-peak hours mean for billing?
How deep is the DeepSeek prompt cache discount?
DeepSeek V4 Flash or V4 Pro?
Related tools
LLM VRAM Calculator
Calculate the VRAM needed to run or fine-tune any LLM at any quantization.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜JSON Formatter
Validate, format, and minify JSON data with syntax highlighting.
Use tool ➜