TechCompare LogoTechCompare

MiniMax-M3 API Pricing & Cost Calculator

MiniMax-M3 is a strong value pick for high-volume utility work - 3x cheaper on output than Gemini Flash, with a real 1M context window. The published 50% batch discount makes it even cheaper at scale.

MiniMax-M3 is MiniMax's frontier offering in the ultra-cheap tier, designed for high-volume classification and routing workloads. A real 1M context window distinguishes it from same-tier peers.

By TechCompare · Updated

Input tokens
30,000
per request
Output tokens
3,000
per request
Volume
1,000 / monthly
Standard API

Calculator

Cost Comparison

Based on 30,000 input tokens (50% cached), 3,000 output tokens, and 1,000 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

How this is calculated

MiniMax-M3 is priced at $0.30 per million input tokens and $1.20 per million output tokens, with an 80% prompt caching discount ($0.06 per million cached reads). A 50% batch discount is also published on OpenRouter ($0.15 per million inputs, $0.60 per million outputs) - one of the few labs to expose batch directly via the OR API.

Verdict

MiniMax-M3 sits at $0.30/M input and $1.20/M output, with the 80% cache discount dropping cached input to $0.06/M and a published 50% batch discount (on OpenRouter) dropping both rows to $0.15/M input and $0.60/M output. The math against Gemini 3.6 Flash is roughly 3x cheaper on output, and the 1M context window matches MiniMax-M3's larger-tier rivals. The cache hit rate decides most of the per-call cost: a high-volume utility pipeline with stable prompts (classification, routing, formatting) sees input collapse to $0.06/M, and where OpenRouter's batch discount applies, the per-call total falls by another half. This is one of the few labs that exposes batch directly via the OR API, which is what unlocks the half-price bulk path.

More API Standalones scenarios

GPT-5.5 Pricing
Single-model gpt-5.5 cost estimate
View details ➜
GPT-5.4 Pricing
Single-model gpt-5.4 cost estimate
View details ➜
Claude Opus 4.8 Pricing
Single-model claude-opus-4.8 cost estimate
View details ➜

Frequently asked questions

How does MiniMax-M3 stack up against GPT-5.6 Luna?
Luna ($0.20/$1.20 list price, and sometimes $0.10/$0.60 during OpenRouter temporary discounts) is cheaper than MiniMax-M3 ($0.30/$1.20), but M3 has a published 50% batch discount and a real 1M context. Pick Luna for the absolute floor on price, M3 for batch-friendlier production.
How much does MiniMax-M3 cost per million output tokens?
MiniMax-M3 charges $1.20 per million output tokens at standard pricing, matching GPT-5.6 Luna's output rate. Input is $0.30 per million (above Luna's $0.20/M). The 80% prompt caching discount brings cached input to $0.06 per million. MiniMax also offers a 50% batch discount on OpenRouter.