i
News
News · 2026-09-22

GPT-6 Sol and Luna halve prices, barely move scores

@neuronium_ai @neuronium_ai

OpenAI has introduced GPT-6 Sol and Luna at half the price of their GPT-5.6 counterparts, but the performance story is less dramatic. Sol costs $2 per million input tokens and $10 per million output tokens; Luna costs $0.10 and $0.50. OpenAI attributes the reduction to better caching and inference, while independent analysis puts both models roughly at GPT-5.6’s intelligence level. The launch looks less like a new capability tier than a bid to make routine model use materially cheaper.

Cover: GPT-6 Sol and Luna halve prices, barely move scores

What shipped

Sol is aimed at recurring, difficult work:

building new features;
code review;
debugging;
data analysis.

Luna targets large volumes of clearly defined operations:

short document summaries;
information extraction;
short answers.

OpenAI says GPT-6 caching can reduce the price of cached input tokens by up to 90%. A new caching dashboard and diagnostic tool are intended to help developers use that discount. Developers can also change the reasoning level and tool availability without resetting the cache.

Terra, previously the cheapest model in the lineup, is no longer available. OpenAI says the new prices bring GPT-6 roughly to the level of cheaper open-weight models.

At launch, both models are available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu subscribers. Free and Go users get Luna through the desktop app, but neither model initially works in the regular chat. In the API, they are available as gpt-6-sol and gpt-6-luna; ChatGPT access is rolling out gradually.

The savings are clearer than the gains

OpenAI’s comparisons focus on price-performance against Claude models rather than on outright leadership.

In OSWorld 2.0, which tests computer control, OpenAI says GPT-6 performs close to Claude Opus 5 at about 80% lower cost. Astra still leads the category.
In AutomationBench, which evaluates workflows using 47 tools, GPT-6 Sol at the maximum reasoning level is reported to outperform Claude Opus 5 at its maximum level. A Sol task costs 9% as much as an Opus task.
Luna improves on its predecessor by 5.4 percentage points while reducing cost by 58%.

The programming results tell a similar story: cheaper execution, limited separation in quality.

On FrontierCode 1.1, which tests whether coding agents produce code that can be integrated into an existing project, GPT-6 Sol scores 49.3% at the maximum reasoning level, at $2.14 per task. Claude Fable 5.1 scores 50.3% but costs $12.83 per task, six times more. Claude Opus 5 scores 53.4% at $4.31 per task using its medium reasoning level, which is its best result on this test.

DeepSWE v1.1 tests difficult software-engineering tasks in real codebases. Here, the gap to the leaders is more visible:

GPT-6 Sol: 68.8% at the maximum reasoning level;
Claude Fable 5: 69.9% at “xhigh” and 69.7% at “max”;
Claude Opus 5: 73.7% at maximum reasoning;
GPT-5.6 Sol: 72.7%.

The reason to choose GPT-6 Sol is therefore not a new quality frontier. It is the ability to approach those results at a much lower cost. At “xhigh,” Sol scores 66.6% for $1 per task. Luna reaches the same result at maximum reasoning for $0.22.

That gap makes the two models unusually close in practical value. Luna matches Sol’s “xhigh” result on DeepSWE while costing 78% less. Raising Sol to maximum reasoning lifts its score to 68.8%—only 2.2 percentage points above Luna.

OpenAI says Luna’s result is comparable to Claude Opus 5 and Claude Fable 5 at their medium reasoning levels. Luna is 93% cheaper than Opus and 96% cheaper than Fable.

Sol maximum68.8% DeepSWE
Luna maximum66.6% DeepSWE

What the benchmarks leave out

The published test set does not include GDPval for knowledge work or Terminal-Bench 4.0 for AI-agent tasks. It also appears not to account for the arrival of Opus 5.5, which may require 40% less spending than Opus 5 while delivering higher performance.

That omission matters because the announcement is built around relative efficiency. A comparison can make GPT-6 look stronger by choosing the tasks and competitors most favorable to its cost profile. I think the more important question is not whether Sol can reach a similar score to an older Claude model for less money, but whether it remains compelling after the comparison set moves.

Artificial Analysis reaches a restrained conclusion. It finds that GPT-6 Sol and Luna halve the cost of individual tasks compared with their predecessors, but remain around GPT-5.6’s level on intelligence measures. Sol gains 2 points in its coding-agent index, while Luna loses 2 points.

On the Artificial Analysis Intelligence Index, GPT-6 Sol climbs from 47 to 48 points over its predecessor while Luna stays at 37. The real difference is the much lower cost per task, as the bottom chart shows. | Image: Artificial Analysis

On the Artificial Analysis Intelligence Index, GPT-6 Sol climbs from 47 to 48 points over its predecessor while Luna stays at 37. The real difference is the much lower cost per task, as the bottom chart shows. | Image: Artificial Analysis

Source: the-decoder.com

Artificial Analysis also reports weaker results on two important knowledge-work benchmarks. GDPval-AA v2.1 tests computer work across 44 professional fields. Sol loses about 100 Elo points and Luna about 75. Manual review attributes most of the regression to lower presentation quality and incomplete results.

Artificial Analysis had previously been criticized for rating Astra too low because of outdated benchmarks. After two tests were updated, Astra returned to first place. It is not yet clear what will happen to the GPT-6 scores.

My read is that OpenAI is selling an operating-cost decision, not a capability decision. For repeated coding, extraction and summarization, a model that is slightly weaker but dramatically cheaper may be the rational choice. But the uneven benchmark results, multiple reasoning levels and omitted tests make the intelligence claim harder to evaluate.

The launch also puts more weight on real daily usage, where benchmarks capture only part of a model’s value. If GPT-6 wins, it may be because teams can afford to run it more often—not because Sol or Luna has decisively moved the frontier.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X