TechCompare LogoTechCompare

MiMo-V2.6-Flash API pricing and cost calculator

The direct half-cached 100K-input, 5K-output workload costs $0.00854 per request. At 100,000 calls, that is $854 at the normal-speed real-time rate.

MiMo-V2.6-Flash charges $0.14 per million uncached input tokens and $0.28 per million output tokens through Xiaomi's overseas pay-as-you-go API. Cache-hit input costs $0.0028 per million tokens. A small per-request price becomes useful when you have many requests or an agent makes several calls for each user action. This page uses Xiaomi's own USD table, checked on October 4, 2026, to show how input reuse, generated output, and request volume change that spending.

By TechCompare · Updated

Input tokens
100,000
per request
Output tokens
5,000
per request
Volume
100 / monthly
Standard API

Calculator

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Normal-speed, real-time provider USD rates, checked October 4, 2026.

Xiaomi: MiMo-V2.6-Flash
$0.85

How this is calculated

For 100,000 input and 5,000 output tokens, uncached input costs $0.014 and output costs $0.0014. The total is $0.0154 per request. If 50,000 input tokens hit the cache, input costs $0.00714 and the total becomes $0.00854. The starting chart shows 100 monthly requests at $0.854, rounded to $0.85. At 100,000 requests, the same half-cached workload costs $854 rather than $1,540 without cache hits. At 10,000 calls, those real-time figures are $85.40 with half the input cached and $154 without cache hits. Reusing input is especially useful when a workflow sends the same instructions or document prefix repeatedly. These estimates assume each request has the stated usage. A user session that needs five model calls doesn't have the cost of one call. Growing conversation histories also change the input length, even when each visible question is short. Track actual cache hits and output tokens across a complete workflow. Xiaomi lists cache writes as free for a limited time, while web search is billed separately. Keep those policies in view when estimating an application that stores long prefixes or searches repeatedly.

Verdict

Flash is worth testing for frequent requests where its results pass your quality checks. Start with extraction, transformations, and coding tasks drawn from the application you plan to run. Check output validity and completed results before assuming that a low price makes it the best default. For tasks that fail those checks, compare the total cost of corrections against a direct Pro call. Cache reuse helps repeated context, while concise outputs can reduce the generated-token portion of the bill.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
GLM-5.3-Flash Pricing
Cost calculator for this model
View details ➜
DeepSeek V4.1 Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

What are Xiaomi's MiMo-V2.6-Flash token prices?
The overseas real-time API charges $0.14 per million uncached input tokens, $0.0028 for cache-hit input, and $0.28 for output. Xiaomi publishes these USD rates directly. The examples use equal billed token counts to explain the math. Your application's actual bill depends on usage, the number of calls, and any separately priced services.
How does Flash pricing compare with Pro?
Flash is about 67.82% cheaper on uncached input and output. Its cache-hit rate is about 22.22% lower. For 100,000 input tokens, half cached, and 5,000 output tokens, Flash costs $0.00854 versus Pro's $0.02628. The price ratio therefore changes with the input, output, and cache mix. Compare successful task costs before choosing which model handles a workload.
How much can output length change Flash spending?
At 100,000 input tokens with half the input cached, 5,000 output tokens give a $0.00854 request. Increasing output to 50,000 tokens raises the total to $0.02114. That becomes $2,114 across 100,000 real-time requests. Keep the output long enough to complete the task, then measure billed reasoning and answer tokens instead of estimating from visible answer length alone.
Does cached input make a Flash request nearly free?
It reduces the input portion substantially, but generated output still costs money. If all 100,000 input tokens hit the cache and the model generates 5,000 output tokens, the total is $0.00168. Longer generated answers raise that output portion. Request a useful output length and verify that you actually receive cache hits rather than applying the cached rate to every prompt.
Does the calculator include Xiaomi search calls?
No. Xiaomi lists overseas web search at $5 per 1,000 uses in addition to token charges. One such use per request adds $500 across 100,000 requests, which can be a meaningful part of a Flash application's budget. Use the provider's recorded search usage and count all model calls needed for each task when estimating total spending.