MiniMax-M3 API Pricing & Cost Calculator

MiniMax-M3 is a strong value pick for high-volume utility work - 3x cheaper on output than Gemini Flash, with a real 1M context window. The published 50% batch discount makes it even cheaper at scale.

MiniMax-M3 is MiniMax's frontier offering in the ultra-cheap tier, designed for high-volume classification and routing workloads. A real 1M context window distinguishes it from same-tier peers.

By TechCompare · Updated

Input tokens
30,000
50% cached
Output tokens
3,000
per request
Volume
1,000 / monthly
Standard API

Calculator

Cost Comparison

Based on 30,000 input tokens (50% cached), 3,000 output tokens, and 1,000 requests.

How this is calculated

MiniMax-M3 is priced at $0.30 per million input tokens and $1.20 per million output tokens, with an 80% prompt caching discount ($0.06 per million cached reads). A 50% batch discount is also published on OpenRouter ($0.15 per million inputs, $0.60 per million outputs) - one of the few labs to expose batch directly via the OR API.

Verdict

MiniMax-M3 is a strong value pick for high-volume utility work - 3x cheaper on output than Gemini Flash, with a real 1M context window. The published 50% batch discount makes it even cheaper at scale.

More API Standalones scenarios

GPT-5.5 Pricing
Single-model gpt-5.5 cost estimate
View details ➜
GPT-5.4 Pricing
Single-model gpt-5.4 cost estimate
View details ➜
Claude Opus 4.8 Pricing
Single-model claude-opus-4.8 cost estimate
View details ➜

Frequently asked questions

How does MiniMax-M3 stack up against GPT-5.6 Luna?
Luna ($0.20/$1.20 list price, and sometimes $0.10/$0.60 during OpenRouter temporary discounts) is cheaper than MiniMax-M3 ($0.30/$1.20), but M3 has a published 50% batch discount and a real 1M context. Pick Luna for the absolute floor on price, M3 for batch-friendlier production.