TechCompare LogoTechCompare

DeepSeek V4.1 Flash API pricing and cost calculator

A half-cached 100K-input, 5K-output request costs $0.01065 off-peak and $0.0213 at peak. The same 100,000 calls cost $1,065 or $2,130 depending on timing.

DeepSeek V4.1 Flash uses time-based pricing on DeepSeek's direct API. Off-peak rates are $0.15 per million uncached input tokens, $0.003 for cache-hit input, and $0.60 for output. Peak rates double each of those figures to $0.30, $0.006, and $1.20. The chart shows both periods using the provider's USD table, checked on October 4, 2026. Request timing can therefore change the bill even when the model, prompt, cache usage, and generated output are identical.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Normal-speed, real-time provider USD rates, checked October 4, 2026.

DeepSeek: V4.1 Flash (off-peak)
$1.07
DeepSeek: V4.1 Flash (peak)
$2.13

How this is calculated

DeepSeek's peak periods are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak, including weekends and those holidays. Keep that UTC schedule in view when scheduling work across time zones. For 100,000 input tokens and 5,000 output tokens without cache hits, the request costs $0.018 off-peak or $0.036 at peak. If half the input hits the cache, the totals fall to $0.01065 and $0.0213. The chart's 100 monthly requests therefore cost $1.065 off-peak, displayed as $1.07, or $2.13 at peak. At 100,000 half-cached requests, the estimates are $1,065 and $2,130. If 30% of that workload runs at peak and the rest off-peak, the blended token cost is $1,384.50. That calculation assumes equal usage per request and no extra attempts. The direct API identifier is deepseek-flash, which the pricing page maps to DeepSeek-V4.1-Flash. Its published context window is 1M tokens and its maximum output is 384K. Those limits allow large workloads, but extra context and generated reasoning still add billable tokens. Both periods use the normal-speed real-time API. The chart shows each period separately so a monthly average can be based on when your requests actually run.

Verdict

Separate interactive traffic from jobs whose timing you can control. Schedule eligible bulk work outside the published peak periods if the delay fits your delivery requirements, then check the billed usage rather than assuming every job stayed in one period. Cache reuse reduces the input portion in either period. For comparisons with another API, use the time period your application actually needs and count all reasoning output, tool turns, and retries required to get an accepted result.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
MiMo-V2.6-Flash Pricing
Cost calculator for this model
View details ➜
GLM-5.3-Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

What are DeepSeek V4.1 Flash's direct USD prices?
Per million tokens, off-peak uncached input is $0.15, cache-hit input is $0.003, and output is $0.60. Peak prices are $0.30, $0.006, and $1.20. Apply the matching period to each request's billed usage. The chart displays both sets of rates rather than selecting a time period from your browser's local clock.
When does DeepSeek charge peak rates?
Peak periods are 01:00 to 04:00 UTC and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Other hours are off-peak, including weekends and those holidays in full. Convert the schedule for your location and account for daylight saving changes. A monthly estimate should use your expected request distribution across those periods.
Which API model name selects V4.1 Flash?
Use deepseek-flash on DeepSeek's API. The provider's pricing page maps that identifier to DeepSeek-V4.1-Flash. It also documents compatibility routing for legacy Flash identifiers. For a fresh integration, use the documented current identifier and check the model version returned by the service rather than assuming a reseller's model ID is accepted on the direct endpoint.
How do I estimate a mix of peak and off-peak traffic?
Calculate each period separately and add the totals. For 100,000 requests with 100,000 input tokens, half cached, and 5,000 output tokens, 70,000 off-peak calls cost $745.50 and 30,000 peak calls cost $639. The combined estimate is $1,384.50. If peak requests have different token usage, use their own input and output averages.
Does moving work off-peak replace prompt caching?
No. Timing and caching affect different parts of the estimate. Moving an identical request off-peak halves all its listed token rates. Cache hits replace the uncached-input rate with the much lower cached-input rate. For a fully cached 100,000-token prompt and 5,000 output tokens, the direct estimate is $0.0033 off-peak or $0.0066 at peak.