OpenAI Says GPT-6 Sol Beats Claude Opus 5 at 9% of the Cost

OpenAI halved API prices for its two cheaper GPT-6 models. The rival scores in its comparisons come from public reports, not a shared test run.

OpenAI released GPT-6 Sol and GPT-6 Luna on September 23, the mid-tier and budget models sitting below the flagship GPT-6 Astra it launched earlier this month. Both are live in the API as gpt-6-sol and gpt-6-luna, and in ChatGPT Work and Codex for paid plans; free and Go users get Luna in the desktop app. OpenAI says both were trained with methods similar to Astra's, and that they cost half as much as the GPT-5.6 models they replace.

The price cut, and what it is measured against

Per million tokens, Sol drops from $4 input and $20 output to $2 and $10. Luna goes from $0.20 and $1.20 to $0.10 and $0.50, which is slightly better than half on output. The baseline matters here: OpenAI compares against GPT-5.6's promotional pricing, so this is a cut from a price that was already discounted. Cached input reads get a 90% discount, and OpenAI says changing reasoning effort or toggling tools mid-conversation no longer throws the cache away, which matters most for long agent sessions.

The benchmark claims

The headline number comes from AutomationBench, a test of business workflows across 47 tools. OpenAI says Sol at its xhigh effort setting scores 33.2% at $0.27 per task, against 26.9% for Claude Opus 5 at max effort at 11.1 times the cost, which is where the 9% figure comes from. OpenAI also claims Sol at max effort lands within 1.1 points of Claude Fable 5 on DeepSWE v1.1 (68.8% against 69.9%) at about 80% lower cost per task, and that Luna at max effort reaches 66.6% on the same test.

Then the fine print. OpenAI ran its own models in its research environment or through the API, and says those runs may differ from production ChatGPT. The competitor numbers came from publicly available reports rather than a shared test run, and where Fable 5.1 had no published score, OpenAI substituted the older Fable 5. The AutomationBench entry for Fable 5.1 leaves out the cost of Opus 5 fallbacks, which OpenAI says happened on about 40% of tasks, so that cost comparison is incomplete by the vendor's own note.

On the cheapest tier, OpenAI says Luna at higher effort matches GPT-5.6 Sol on its factuality evaluation at about a hundredth of the cost. That evaluation uses conversations where users had already flagged a factual error, which OpenAI says is not representative of typical use.

What the release does not say

There are no parameter counts, context window figures or latency numbers, and no independent test yet. Sol also still trails Fable 5 on DeepSWE, by OpenAI's own figures; the pitch is price per task, not the top score. Neither model is in ChatGPT's standard Chat mode yet.

Sources