i
News
News · 2026-09-29

OpenAI prices GPT-6.1 Sol five times below Astra and sells speed

@neuronium_ai @neuronium_ai

OpenAI has added GPT-6.1 Sol, a lower-cost model it says has narrowed the gap with GPT-6 Astra, and introduced Ultrafast, a paid inference mode that can generate up to 300 tokens per second. The two launches put different costs on the same engineering trade-off: Sol makes repeated agent work cheaper, while Ultrafast charges a premium to reduce waiting. For teams building AI into workflows, the choice is increasingly not just which model to use, but how much speed each step is worth.

Cover: OpenAI prices GPT-6.1 Sol five times below Astra and sells speed

Sol’s price is the story

GPT-6.1 Sol keeps GPT-6 Sol’s standard prices: $2 per million input tokens and $10 per million output tokens. Cached input tokens fall from $0.20 to $0.10 per million, a 50% cut. Moving to the new model therefore does not raise the basic token budget.

Against GPT-6 Astra, Sol’s ordinary input and output tokens cost five times less; cached inputs cost ten times less. At today’s discounted prices, GPT-6.1 Sol is also half the cost of GPT-5.6 Sol. OpenAI says that discount will remain in place at least through November 21, 2026.

The cache reduction matters most for systems that reuse long prompts: coding agents, research systems and business-process automation. When an agent makes hundreds or thousands of calls, repeatedly paying for the same instructions and reference material can add up.

OpenAI says Sol has also closed much of the performance gap with Astra in its own evaluations:

On DeepSWE v1.1, a benchmark for long software-engineering tasks in real codebases, Sol matched Astra at roughly one-fifth the cost. It beat GPT-6 Sol’s best result by 6.4 percentage points.
On GDP.pdf, a benchmark for complex professional documents, OpenAI says Sol scored above Claude Opus 5.5 with fallback models, at less than half the cost per task. It also approached Astra’s result at roughly one-fifth the cost.
On AutomationBench, which tests end-to-end tasks across 47 tools, Sol beat Opus 5.5 by 2.2 percentage points at medium reasoning, for roughly one-third the cost. Its score was 4.8 percentage points above GPT-6 Sol under the same conditions.
On the offline OSWorld 2.0 computer-use benchmark, Sol at maximum reasoning trailed Astra by 2.1 percentage points, with each task costing about one-seventh as much.

These are OpenAI evaluations, not independent tests on production systems. The company says its results come from research settings or API use and may differ from ChatGPT in practice. That distinction matters: the benchmarks support a price-performance case, but they do not establish how Sol will perform inside a company’s own workflows.

Ultrafast sells less waiting

Ultrafast is a separate inference mode, not a cheaper way to run the same workload. OpenAI says it can reach 300 tokens per second, increase speed by up to 8 times in Codex and by up to 6 times through the API. API use costs six times the standard rate.

For comparison, OpenAI’s existing Fast mode costs twice the standard rate for GPT-6 models. That makes Ultrafast three times the price of Fast and six times the price of Standard for the same base model.

The company has published the multiplier, but not separate Ultrafast prices for each model. Applying it to Astra’s standard rates implies $60 per million input tokens and $300 per million output tokens, compared with $10 and $50 at Standard. For GPT-6.1 Sol, the same calculation implies $12 per million input tokens and $60 per million output tokens; cached input would be $0.60 per million if the multiplier applies to that price too. These are estimates based on the announced multiplier, not published model-specific rates.

Standard1× price
Fast2× price

That pricing makes Ultrafast a poor default for background agents or document processing, where throughput matters more than how quickly a person sees each response. It is aimed instead at interactive work: coding, support, financial analysis and incident response.

OpenAI has tested the approach before. In August, it introduced GPT-5.6 Sol Ultrafast on Cerebras hardware, with up to 750 output tokens per second and speeds up to 14 times faster than Standard. Access was limited. DevDay turns that experiment into a broader product offer, but the new GPT-6 models have lower maximum stated speeds than that earlier system.

A choice of cost, capability and latency

GPT-6.1 Sol is available through the API as gpt-6.1-sol, and to ChatGPT Work and Codex users on Plus, Pro, Business, Enterprise and Edu plans. It is not yet available in the regular ChatGPT product. GPT-6 Astra Ultrafast is available through the API and to Pro 500 and Enterprise users in ChatGPT Work and Codex; Sol Ultrafast is expected in the coming days.

I think the more interesting question is how often a workflow truly needs Ultrafast. OpenAI’s pricing makes the trade-off explicit, but the announcement does not quantify how much faster responses translate into better outcomes for the applications it names. A sixfold price premium is easy to calculate; the value of saving time depends on the task.

For enterprise teams, the two launches point in opposite directions: run routine agent work on a model that costs less, and pay for speed only where waiting has a real cost. That gives architects a finer control over inference budgets, while leaving the hard part where it has always been: deciding which parts of a workflow deserve the premium.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X