GPT-6 Luna vs GPT-6.1 Sol: API cost versus reasoning
Luna cuts base input and output costs by 95%. Use Sol where your evaluations show that better completion quality pays for its higher token bill.
Luna's base input and output rates are 20 times lower. Compare the cost of a successful result.
GPT-6 Luna's base input and output rates are 20 times lower than GPT-6.1 Sol's. That's a large saving for repeated requests, but the useful comparison is the cost of getting an acceptable result. A cheap call that needs several corrections can erase part of its advantage. This article compares equal token workloads first, then explains how to test whether Sol's reasoning is worth the extra spend on your own tasks.
By TechCompare · Updated
Cost Comparison
Based on 100,000 input tokens (50% cached), 5,000 output tokens, and 100 requests.Prices are fetched live from OpenRouter and may include temporary promotional discounts not accounted for in our article and comparison figures.
Side-by-side specs
| Spec | GPT-6 Luna | GPT-6.1 Sol |
|---|---|---|
| Standard input per 1M tokens | $0.10 up to 272K input (better on this spec) | $2.00 up to 272K input |
| Standard output per 1M tokens | $0.50 up to 272K input (better on this spec) | $10.00 up to 272K input |
| Cache reads per 1M tokens | $0.01 up to 272K input (better on this spec) | $0.10 up to 272K input |
| 20K input + 2K output, no cache hits | $0.003 (better on this spec) | $0.06 |
| 20K input + 2K output, 50% cached input | $0.0021 (better on this spec) | $0.041 |
| 100K input + 5K output, 50% cached input | $0.008 (better on this spec) | $0.155 |
| 300K input + 5K output, no cache hits | $0.06375 (better on this spec) | $1.275 |
| Input / output above 272K input, per 1M | $0.20 / $0.75 | $4.00 / $15.00 |
| Context window / maximum output | 1,050,000 / 128,000 tokens | 1,050,000 / 128,000 tokens |
| Provider Batch API discount | 50% | 50% |
How they differ
Standard input costs $0.10 per million tokens on Luna and $2 on Sol. Output costs $0.50 and $10, while cache reads cost $0.01 and $0.10. A request with 20,000 input and 2,000 output tokens costs $0.003 on Luna and $0.06 on Sol without cache hits. Reading half the input from cache changes those totals to $0.0021 and $0.041. The gap becomes smaller because Sol discounts cached input more heavily. Both models have a 1,050,000-token context window and a full-request surcharge above 272,000 input tokens. At 300,000 input and 5,000 output tokens without cache hits, the estimates are $0.06375 and $1.275. These figures exclude cache writes and tool fees. The embedded chart uses live OpenRouter rates for its displayed workload.
Verdict
Test a representative task set before choosing one model for all traffic. Start with tasks that have clear success checks, such as valid extracted fields, passing code checks, or a correct calculation. Record total billed usage, corrections, and review time for each accepted result. Route routine work to Luna if it meets your standard, then use Sol for cases where measured improvements justify the cost. Price alone doesn't establish which model reasons better on a particular problem.
Which should you pick?
Choose GPT-6 Luna
Choose Luna when focused requests pass your quality checks and you expect enough volume for small per-call savings to matter. Try it on classification, extraction, and short transformations using examples from your application. Keep outputs concise where the task allows it, and track failures so a low token bill doesn't hide extra manual work or repeated attempts.
Official Luna model referenceChoose GPT-6.1 Sol
Choose Sol when your task set shows that it resolves difficult cases with fewer corrections or less review. Examples to test include changes across several code files, ambiguous instructions, and analysis that must satisfy several constraints. Define that improvement before spending more. A higher reasoning setting can also change output usage, so compare completed tasks using billed tokens rather than equal settings alone.
Official Sol model reference