DeepSeek V4.1 Flash API pricing and cost calculator
A half-cached 100K-input, 5K-output request costs $0.01065 off-peak and $0.0213 at peak. The same 100,000 calls cost $1,065 or $2,130 depending on timing.
DeepSeek V4.1 Flash uses time-based pricing on DeepSeek's direct API. Off-peak rates are $0.15 per million uncached input tokens, $0.003 for cache-hit input, and $0.60 for output. Peak rates double each of those figures to $0.30, $0.006, and $1.20. The chart shows both periods using the provider's USD table, checked on October 4, 2026. Request timing can therefore change the bill even when the model, prompt, cache usage, and generated output are identical.
By TechCompare · Updated
Calculator
Cost Comparison
Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Normal-speed, real-time provider USD rates, checked October 4, 2026.
How this is calculated
DeepSeek's peak periods are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak, including weekends and those holidays. Keep that UTC schedule in view when scheduling work across time zones. For 100,000 input tokens and 5,000 output tokens without cache hits, the request costs $0.018 off-peak or $0.036 at peak. If half the input hits the cache, the totals fall to $0.01065 and $0.0213. The chart's 100 monthly requests therefore cost $1.065 off-peak, displayed as $1.07, or $2.13 at peak. At 100,000 half-cached requests, the estimates are $1,065 and $2,130. If 30% of that workload runs at peak and the rest off-peak, the blended token cost is $1,384.50. That calculation assumes equal usage per request and no extra attempts. The direct API identifier is deepseek-flash, which the pricing page maps to DeepSeek-V4.1-Flash. Its published context window is 1M tokens and its maximum output is 384K. Those limits allow large workloads, but extra context and generated reasoning still add billable tokens. Both periods use the normal-speed real-time API. The chart shows each period separately so a monthly average can be based on when your requests actually run.
Verdict
Separate interactive traffic from jobs whose timing you can control. Schedule eligible bulk work outside the published peak periods if the delay fits your delivery requirements, then check the billed usage rather than assuming every job stayed in one period. Cache reuse reduces the input portion in either period. For comparisons with another API, use the time period your application actually needs and count all reasoning output, tool turns, and retries required to get an accepted result.
More API Standalones scenarios
Related guides
Frequently asked questions
What are DeepSeek V4.1 Flash's direct USD prices?
When does DeepSeek charge peak rates?
Which API model name selects V4.1 Flash?
How do I estimate a mix of peak and off-peak traffic?
Does moving work off-peak replace prompt caching?
Related tools
LLM VRAM Calculator
Calculate the VRAM needed to run or fine-tune any LLM at any quantization.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜JSON Formatter
Validate, format, and minify JSON data with readable output and error detection.
Use tool ➜