TechCompare LogoTechCompare

GPT-6 Luna API pricing and cost calculator

Luna's 20K-input, 2K-output example costs $0.003 uncached or $0.0021 with half the input cached. At 100,000 monthly requests, those token totals are $300 and $210.

GPT-6 Luna's standard base rates are $0.10 per million input tokens and $0.50 per million output tokens, with cache reads at $0.01 per million. Those prices make it worth evaluating for repeated classification, extraction, and routing work. At high request volumes, small per-call differences add up. The relevant questions are whether it completes the task reliably and how much spending remains once you include volume, retries, tools, and cache creation.

By TechCompare · Updated

Input tokens
20,000
per request
Output tokens
2,000
per request
Volume
100,000 / monthly
Standard API

Calculator

Cost Comparison

Based on 20,000 input tokens (50% cached), 2,000 output tokens, and 100,000 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

How this is calculated

This calculator starts with 20,000 input tokens, 2,000 output tokens, half the input read from cache, and 100,000 monthly requests. Uncached input costs $0.001, cache-read input costs $0.0001, and output costs $0.001, giving $0.0021 per request or $210 per month. Without cache hits, the same workload is $0.003 per call or $300 monthly. Above 272,000 input tokens, Luna applies double input and cache prices and 1.5 times output prices to the whole request. At 300,000 input and 5,000 output tokens, that is $0.06375 without cache hits or $0.03675 with half the input cached. Cache writes and tools are separate. The embedded chart loads live OpenRouter provider prices, while these examples use OpenAI's standard rates.

Verdict

Evaluate Luna directly on the small tasks that dominate your application's request count. Keep outputs short when the task only needs a label or a few extracted fields, and track failures rather than assuming cheap calls make retries harmless. GPT-6.1 Sol is a useful next comparison when you need to decide which tasks deserve a larger reasoning budget. A standalone Luna estimate is already useful for seeing whether volume or tools dominate the bill.

More API Standalones scenarios

MiMo-V2.6-Pro Pricing
Cost calculator for this model
View details ➜
MiMo-V2.6-Flash Pricing
Cost calculator for this model
View details ➜
GLM-5.3-Flash Pricing
Cost calculator for this model
View details ➜

Frequently asked questions

What would one million example Luna requests cost?
At 20,000 input and 2,000 output tokens per call, one million requests cost $3,000 without cache hits or $2,100 with half the input read from cache at standard rates. These are model token totals. Writes, tools, retries, and any provider or processing premiums are additional. Even a small per-request difference matters when it is repeated a million times.
Does a cheap model mean the whole agent is cheap?
No. An agent can make several calls per task, accumulate long prompts, and use separately billed tools. Its full cost depends on all of those steps. Measure how many calls are needed for an accepted result and whether a shorter output format reduces spending. Cheap tokens help, but they don't remove the need to budget the workflow.
How much does Luna prompt caching save?
At base rates, cache reads are $0.01 per million tokens rather than $0.10 for uncached input, a 90% read discount. Cache writes cost $0.125 per million, so repeated reads need to cover that creation cost. Only actual cache-hit tokens qualify. The calculator's percentage models reads and leaves write charges outside the estimate.
Can GPT-6 Luna use large prompts at its base price?
Not across its entire context window. Input above 272,000 tokens moves the full request to $0.20 input, $0.02 cache reads, and $0.75 output per million tokens. The cache-write rate doubles to $0.25 too. Check that boundary when processing large documents, even if the model can accept the prompt.
What does Luna cost through the Batch API?
OpenAI's Batch API halves applicable token rates. The 20K-input, 2K-output example becomes $0.0015 uncached or $0.00105 with half the input cached before write costs and tools. At 100,000 requests, that is $150 or $105. Use those figures for an asynchronous provider workflow rather than an assumed discount on interactive OpenRouter requests.
What should GPT-6 Luna be compared with?
GPT-6.1 Sol is useful for deciding when a task needs a larger model budget: its base input and output rates are twenty times Luna's. Claude Haiku 4.5 is another candidate when comparing providers for frequent focused requests. Neither price ratio proves which completes your task better. Hold the success criteria constant and compare total spending per accepted result.