MiMo-V2.6-Flash API pricing and cost calculator
The direct half-cached 100K-input, 5K-output workload costs $0.00854 per request. At 100,000 calls, that is $854 at the normal-speed real-time rate.
MiMo-V2.6-Flash charges $0.14 per million uncached input tokens and $0.28 per million output tokens through Xiaomi's overseas pay-as-you-go API. Cache-hit input costs $0.0028 per million tokens. A small per-request price becomes useful when you have many requests or an agent makes several calls for each user action. This page uses Xiaomi's own USD table, checked on October 4, 2026, to show how input reuse, generated output, and request volume change that spending.
By TechCompare · Updated
Calculator
Cost Comparison
Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Normal-speed, real-time provider USD rates, checked October 4, 2026.
How this is calculated
For 100,000 input and 5,000 output tokens, uncached input costs $0.014 and output costs $0.0014. The total is $0.0154 per request. If 50,000 input tokens hit the cache, input costs $0.00714 and the total becomes $0.00854. The starting chart shows 100 monthly requests at $0.854, rounded to $0.85. At 100,000 requests, the same half-cached workload costs $854 rather than $1,540 without cache hits. At 10,000 calls, those real-time figures are $85.40 with half the input cached and $154 without cache hits. Reusing input is especially useful when a workflow sends the same instructions or document prefix repeatedly. These estimates assume each request has the stated usage. A user session that needs five model calls doesn't have the cost of one call. Growing conversation histories also change the input length, even when each visible question is short. Track actual cache hits and output tokens across a complete workflow. Xiaomi lists cache writes as free for a limited time, while web search is billed separately. Keep those policies in view when estimating an application that stores long prefixes or searches repeatedly.
Verdict
Flash is worth testing for frequent requests where its results pass your quality checks. Start with extraction, transformations, and coding tasks drawn from the application you plan to run. Check output validity and completed results before assuming that a low price makes it the best default. For tasks that fail those checks, compare the total cost of corrections against a direct Pro call. Cache reuse helps repeated context, while concise outputs can reduce the generated-token portion of the bill.
More API Standalones scenarios
Related guides
Frequently asked questions
What are Xiaomi's MiMo-V2.6-Flash token prices?
How does Flash pricing compare with Pro?
How much can output length change Flash spending?
Does cached input make a Flash request nearly free?
Does the calculator include Xiaomi search calls?
Related tools
LLM VRAM Calculator
Calculate the VRAM needed to run or fine-tune any LLM at any quantization.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜JSON Formatter
Validate, format, and minify JSON data with readable output and error detection.
Use tool ➜