TechCompare LogoTechCompare

GPT-6 Luna vs GPT-6.1 Sol: API cost versus reasoning

Luna cuts base input and output costs by 95%. Use Sol where your evaluations show that better completion quality pays for its higher token bill.

Luna's base input and output rates are 20 times lower. Compare the cost of a successful result.

GPT-6 Luna's base input and output rates are 20 times lower than GPT-6.1 Sol's. That's a large saving for repeated requests, but the useful comparison is the cost of getting an acceptable result. A cheap call that needs several corrections can erase part of its advantage. This article compares equal token workloads first, then explains how to test whether Sol's reasoning is worth the extra spend on your own tasks.

By TechCompare · Updated

Cost Comparison

Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.

Option A
GPT-6 Luna
Wins 7 of 10 compared specs
Option B
GPT-6.1 Sol
Wins 0 of 10 compared specs

Side-by-side specs

SpecGPT-6 LunaGPT-6.1 Sol
Standard input per 1M tokens
$0.10 up to 272K input (better on this spec)
$2.00 up to 272K input
Standard output per 1M tokens
$0.50 up to 272K input (better on this spec)
$10.00 up to 272K input
Cache reads per 1M tokens
$0.01 up to 272K input (better on this spec)
$0.10 up to 272K input
20K input + 2K output, no cache hits
$0.003 (better on this spec)
$0.06
20K input + 2K output, 50% cached input
$0.0021 (better on this spec)
$0.041
100K input + 5K output, 50% cached input
$0.008 (better on this spec)
$0.155
300K input + 5K output, no cache hits
$0.06375 (better on this spec)
$1.275
Input / output above 272K input, per 1M
$0.20 / $0.75
$4.00 / $15.00
Context window / maximum output
1,050,000 / 128,000 tokens
1,050,000 / 128,000 tokens
Provider Batch API discount
50%
50%

How they differ

Standard input costs $0.10 per million tokens on Luna and $2 on Sol. Output costs $0.50 and $10, while cache reads cost $0.01 and $0.10. A request with 20,000 input and 2,000 output tokens costs $0.003 on Luna and $0.06 on Sol without cache hits. Reading half the input from cache changes those totals to $0.0021 and $0.041. The gap becomes smaller because Sol discounts cached input more heavily. Both models have a 1,050,000-token context window and a full-request surcharge above 272,000 input tokens. At 300,000 input and 5,000 output tokens without cache hits, the estimates are $0.06375 and $1.275. These figures exclude cache writes and tool fees. The embedded chart uses live OpenRouter rates for its displayed workload.

Verdict

Test a representative task set before choosing one model for all traffic. Start with tasks that have clear success checks, such as valid extracted fields, passing code checks, or a correct calculation. Record total billed usage, corrections, and review time for each accepted result. Route routine work to Luna if it meets your standard, then use Sol for cases where measured improvements justify the cost. Price alone doesn't establish which model reasons better on a particular problem.

Which should you pick?

Choose GPT-6 Luna

Choose Luna when focused requests pass your quality checks and you expect enough volume for small per-call savings to matter. Try it on classification, extraction, and short transformations using examples from your application. Keep outputs concise where the task allows it, and track failures so a low token bill doesn't hide extra manual work or repeated attempts.

Official Luna model reference

Choose GPT-6.1 Sol

Choose Sol when your task set shows that it resolves difficult cases with fewer corrections or less review. Examples to test include changes across several code files, ambiguous instructions, and analysis that must satisfy several constraints. Define that improvement before spending more. A higher reasoning setting can also change output usage, so compare completed tasks using billed tokens rather than equal settings alone.

Official Sol model reference

Related comparisons

MiMo-V2.6-Pro vs MiMo-V2.6-Flash
Compare Xiaomi's normal-speed real-time rates, cache savings, and the cost of successful tasks.
Read comparison ➜
GLM-5.3-Flash vs DeepSeek V4.1 Flash
Cache hits can change the cheaper option off-peak. Compare Z.ai and DeepSeek's direct rates.
Read comparison ➜
MiMo-V2.6-Flash vs DeepSeek V4.1 Flash
Compare Xiaomi's flat real-time rates with DeepSeek's peak and off-peak token prices.
Read comparison ➜
GPT-6 Luna vs Claude Haiku 4.5
Luna's base token rates are ten times lower. Task quality and integration decide whether switching pays off.
Read comparison ➜
GPT-6.1 Sol vs Claude Opus 5.5
Sol costs half as much at base rates, but its long-prompt surcharge narrows the gap.
Read comparison ➜
GPT-6.1 Sol vs Claude Sonnet 5.5
Both charge $2 input and $10 output per million tokens. Cache reuse and prompt length separate them.
Read comparison ➜

Frequently asked questions

What would 100,000 requests cost on Luna and Sol?
At 20,000 input and 2,000 output tokens per call without cache hits, the token estimates are $300 on Luna and $6,000 on Sol. If half the input is read from cache, they become $210 and $4,100. For the same uncached workload, each provider's Batch API halves the estimate to $150 and $3,000. Cache writes, tools, and other processing charges are additional.
Is Sol worth paying more for reasoning?
That depends on the work it completes successfully in your evaluations. Compare both models on the same tasks and score correctness, instruction following, extra calls, and review time. Both support reasoning, so the distinction isn't simply whether a model can think. Include billed reasoning output in your budget and pay for Sol where its measured improvement matters to the result.
Does Luna stay 20 times cheaper with caching?
The base input and output rates have a twentyfold gap, but cache reads have a tenfold gap. With 20,000 input tokens, half cached, and 2,000 output tokens, Sol's $0.041 estimate is about 19.5 times Luna's $0.0021. The ratio depends on the input, output, and cache mix. Creating a cache also costs money, so include writes when estimating a fresh or frequently changing prompt.
Can I send easy tasks to Luna and hard tasks to Sol?
Yes, if you have a way to identify cases that need another model. For example, routing 90% of the half-cached 20K-input, 2K-output workload directly to Luna and 10% directly to Sol costs $599 over 100,000 calls. That assumes one call per task. A workflow that tries Luna first and then calls Sol needs to include both calls for escalated tasks.