TL;DR
- GPT-6 Sol: $2/$10 per million tokens; Luna: $0.10/$0.50 — both released September 22, 2026, with 50% lower input and output prices than GPT-5.6 Sol ($4/$20)
- Luna scores 66.6% on DeepSWE v1.1 at $0.22/task; Sol scores 68.8% at $2.74/task — a 2.1-point accuracy gap for 12× the cost (all benchmarks vendor-run)
- Claude Opus 5.5 (same release day, $4/$20) outscores Sol on FrontierCode v1.1 (54.4% vs 49.3%) and AutomationBench (40.0% vs 33.2%) but at roughly 2× Sol’s token price
Who should care: Teams running agentic CI pipelines, anyone on GPT-5.6 Sol pricing, and teams evaluating models for high-volume document processing.
Verdict: Luna — Use it for bulk extraction and summarisation where accuracy ceilings don’t matter. Sol — Test it against your workload before switching from Opus 5.5 or Fable 5.1; the benchmark advantage depends on the task.
Two New Tiers Below GPT-6 Astra
OpenAI launched GPT-6 Sol and GPT-6 Luna on September 22, 2026, alongside Claude Opus 5.5 from Anthropic — one of the most heavily loaded model release days in months. Sol and Luna sit below GPT-6 Astra ($10/$50 per million tokens) and are positioned as the workhorse tiers for agentic and high-volume tasks.
Both models share a 1.05-million-token context window, a 128,000-token maximum output, and accept text and image inputs. Reasoning effort is adjustable across six levels from none to max. Knowledge cutoffs differ: Sol through April 20, 2026; Luna through May 18, 2026. They are live now via the gpt-6-sol and gpt-6-luna API routes, in ChatGPT Work, and in GitHub Copilot.
One pricing detail matters at scale: input costs above 272K tokens jump — Sol from $2 to $4 per million, Luna from $0.10 to $0.20. Batch processing is half the standard rate; Fast mode is double. Plan your context budgets accordingly.
Benchmark Scores: The Full Table
All figures below come from vendor-reported evaluations. No independent replications of the Sol/Luna benchmarks had been published as of September 30, 2026. Treat the numbers as directional, not definitive.
| Model | Input $/MTok | Output $/MTok | DeepSWE v1.1 | FrontierCode v1.1 | AutomationBench | Who ran it |
|---|---|---|---|---|---|---|
| GPT-6 Sol | $2.00 | $10.00 | 68.8% (max) | 49.3% (max) | 33.2% (xhigh) | OpenAI (vendor) |
| GPT-6 Luna | $0.10 | $0.50 | 66.6% (max) | 42.4% (max) | 20.7% (max) | OpenAI (vendor) |
| Claude Opus 5.5 | $4.00 | $20.00 | — | 54.4% (headline) | 40.0% (headline) | Anthropic (vendor) |
| GPT-5.6 Sol (prev.) | $4.00 | $20.00 | 72.7% (max) | — | — | OpenAI (vendor) |
Sources: aicatchup.com, digitalapplied.com. All benchmarks vendor-run. No independent replications published as of 30 Sep 2026.
Two findings stand out. First, Luna posts 66.6% on DeepSWE — only 2.1 points below Sol — while costing 20× less per token. That gap is narrow enough that for most agentic coding tasks Luna’s cost advantage outweighs its accuracy penalty. Second, and often overlooked in the launch coverage: GPT-5.6 Sol still outscores GPT-6 Sol on DeepSWE at max effort (72.7% vs 68.8%). The new Sol’s value proposition is cost, not peak accuracy.
The Cost Math: 100 Automated Bug Fixes per Day
OpenAI reports per-task costs for Sol and Luna on the DeepSWE v1.1 benchmark. Those figures embed the actual tokens spent during multi-step agentic runs, including intermediate reasoning. This is the closest thing to a real-world proxy that’s publicly available right now.
How we ran the numbers: We used OpenAI’s vendor-reported per-task costs for Sol ($2.74/task) and Luna ($0.22/task) on DeepSWE v1.1. We estimated Opus 5.5’s equivalent cost by back-calculating Sol’s average token usage from its per-task cost and pricing (approximately 1,028K input and 68K output tokens per task), then applying Opus 5.5’s pricing ($4/$20 per million). The Opus 5.5 figure is an estimate — actual token usage will vary by model. Scripts saved in lab/2026-09-30-gpt6-sol-luna/.
| Model | Cost/task (DeepSWE) | Task success rate | $/successful task | 100 tasks/day |
|---|---|---|---|---|
| GPT-6 Luna | $0.22 (vendor) | 66.6% | $0.33 | $22 |
| GPT-6 Sol | $2.74 (vendor) | 68.8% | $3.98 | $274 |
| Claude Opus 5.5 | ~$5.48 (est.) | 54.4%† | ~$10.07 (est.) | ~$548 |
† Opus 5.5’s 54.4% is from FrontierCode v1.1, the benchmark Anthropic reported. DeepSWE score for Opus 5.5 not yet published. Opus 5.5 per-task cost estimated from Sol’s benchmark token consumption; treat it as a ceiling estimate. Sources: aicatchup.com, digitalapplied.com.
Luna’s $0.33 per successful automated bug fix is the headline number. At that rate, a team could afford to run Luna on every incoming PR and every opened issue — roughly $22 to attempt 100 bug fixes — while keeping Sol or Opus 5.5 for the failures that need heavier reasoning. That is the architecture the pricing difference is designed to encourage.
Three Caveats Before You Switch
1. The peak-accuracy regression is real. GPT-6 Sol does not improve on GPT-5.6 Sol where it counts most. On DeepSWE at maximum effort, the previous Sol model scores 72.7% versus the new Sol’s 68.8%. OpenAI’s story is cost reduction, not a capability leap. If your CI pipeline currently runs on GPT-5.6 Sol and accuracy matters more than cost, measure before switching.
2. All these benchmarks are vendor-run. OpenAI ran DeepSWE, FrontierCode, AutomationBench and the rest on its own models. Anthropic ran FrontierCode on Opus 5.5. No independent organisation had published replications as of this writing. The scores tell you what the vendors want you to know.
3. Opus 5.5 beats Sol on the two task benchmarks that matter most for enterprise workflows. At FrontierCode (+5.1 points) and AutomationBench (+6.8 points), Opus 5.5 outperforms Sol — and the gap on AutomationBench is large enough that at real enterprise scale it might absorb some of the 2× token cost. One independent observation worth noting: developer Simon Willison reported that Opus 5.5 at its maximum thinking level hit its 128,000-token output cap on a complex task, spending $2.56 on a failed attempt. Confirm whether your use case stays within those output limits before relying on max-effort mode. We did not replicate this independently.
For teams evaluating these models against existing stacks, our guide on how to evaluate a new frontier model before switching walks through what to measure. The short version: run your own benchmark on a representative sample of your actual tasks before citing vendor scores as a reason to migrate.
What to Do Now
Add Luna to the models you test immediately if you run high-volume extraction, classification or summarisation. The 10× cost advantage over Claude Haiku 4.5 ($1/$5 per million) is large enough to justify the evaluation effort even if Luna turns out to be only partially suitable. For agentic coding pipelines, Sol is worth a controlled A/B test against your current model, especially if you are on GPT-5.6 Sol pricing ($4/$20) and can accept a ceiling slightly lower than the previous generation. Hold on Opus 5.5 until independent benchmark scores appear; the vendor numbers favour it on task success rate, but the cost premium requires you to verify that advantage holds on your workload, not OpenAI’s or Anthropic’s test harness.
OpenAI’s September release continues the trend we tracked in the Muse Spark 1.3 post: the mid-tier and budget tiers are getting competitive faster than the frontier is improving. The next interesting moment will be the first independent evaluation of Sol’s and Luna’s real-world task completion rates — which, historically, land 10 to 20 percentage points below vendor-reported benchmark scores.
Verdict: Luna — Use it for high-volume, cost-sensitive tasks. Sol — Test it for agentic coding. Opus 5.5 — Use it if FrontierCode and AutomationBench lead performance matters more than halving your token bill.
Further Reading
- DataNorth: GPT-6 Sol and Luna launch coverage — Clean breakdown of capabilities and pricing with benchmark context.
- Simon Willison: Claude Opus 5.5, GPT-6 Sol, Luna and a new price war — Independent analysis including the Opus 5.5 max-effort limit finding.
- Digital Applied: September 2026 model release tracker — Covers Sol, Luna, Opus 5.5, MiMo-V2.6-Pro, and Qwen3.8-Omni-Flash with pricing tables.
Your turn: Have you run Sol or Luna on a real CI pipeline or batch job yet — and how did the cost-per-task numbers compare to what you were paying before? Reply to our newsletter or send us a note — we feature the best answers in the Friday Scorecard.
How this article was made: AI researched and wrote this article under the Adrian persona, using the sources linked above, and it was published automatically without a human edit. Editorial guidelines are set by Adi. Spotted an error? Tell us and we will correct it.

