MiMo-V2.6-Pro API pricing and cost calculator
Xiaomi's direct half-cached 100K-input, 5K-output example costs $0.02628 per request, or $262.80 for 10,000 calls.
MiMo-V2.6-Pro costs $0.435 per million uncached input tokens and $0.87 per million output tokens on Xiaomi's overseas pay-as-you-go API. Input that hits the prompt cache costs $0.0036 per million tokens. These are Xiaomi's published USD rates, checked on October 4, 2026. The chart uses the same direct provider prices, so its starting estimate matches the calculations below. For an agent or document workflow, the useful budget includes every call needed to finish the task, including reasoning output, tool responses, and retries.
By TechCompare · Updated
Calculator
Cost Comparison
Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Normal-speed, real-time provider USD rates, checked October 4, 2026.
How this is calculated
A request with 100,000 input tokens and 5,000 output tokens costs $0.0435 for uncached input plus $0.00435 for output, totaling $0.04785. If half the input hits the cache, input falls to $0.02193 and the request total becomes $0.02628. At 100 monthly requests, the starting chart is $2.628, displayed as $2.63. At 10,000 identical requests, that is $262.80 with half the input cached or $478.50 without cache hits. A fully cached version of that prompt costs $0.00471 per request, including the same output. This makes reusable context valuable, while generated output becomes a larger share of the total as the input bill falls. Xiaomi lists cache writes as free for a limited time. That policy should be checked again before building a long-term budget around it. A cache-hit percentage describes input tokens actually reused, not a discount applied to the entire request. Output remains billed at the real-time rate.
Verdict
Evaluate Pro where complex tasks might justify spending more than the Flash rate. Use the same prompts and success criteria, then record accepted results, output usage, repeated tool calls, and time spent reviewing them. A lower failure rate can matter more than the price of one attempt, but that improvement needs to show up in your application. Stable prompt prefixes can keep input spending low. Keep generated explanations and repeated tool results within useful limits because cache savings don't remove the output bill.
More API Standalones scenarios
Related guides
Frequently asked questions
What does MiMo-V2.6-Pro cost directly from Xiaomi?
How much does MiMo Pro save through prompt caching?
What if all the input is read from cache?
Are Xiaomi Token Plan credits the same as API pricing?
What charges are outside this estimate?
Related tools
LLM VRAM Calculator
Calculate the VRAM needed to run or fine-tune any LLM at any quantization.
Use tool ➜Power Cost Estimator
Estimate annual electricity costs for your PC, Server, or TV.
Use tool ➜Data Transfer Calculator
Estimate transfer times for files over USB, WiFi, Ethernet, and more.
Use tool ➜JSON Formatter
Validate, format, and minify JSON data with readable output and error detection.
Use tool ➜