i
News
News · 2026-09-24

AI benchmark prices are collapsing, but the best models still cost more

@neuronium_ai @neuronium_ai

Epoch AI says the market price of reaching a fixed AI benchmark result is falling faster than for any earlier technology capable of reshaping industries. Its estimate is striking: reproducing OpenAI o3’s 75% score on GPQA Diamond went from 30 cents per question to four hundredths of a cent in 18 months. But a second MIT study finds that algorithmic efficiency explains only part of the decline—and that the best available model can still cost more per request.

Cover: AI benchmark prices are collapsing, but the best models still cost more

The same score, at 1/725 of the price

Epoch AI’s comparison uses five benchmarks covering mathematics, natural sciences and logic puzzles. On GPQA Diamond, a doctoral-level science test, OpenAI o3 reached 75% in early 2025 at an estimated cost of 30 cents per question. A model from the GPT-5.6 family reached the same score 18 months later for four hundredths of a cent.

That is 1/725 of the original price. At the same rate, a €50,000 car would cost less than €70.

OpenAI released the cheaper GPT-6 Sol and Luna models just days before the comparison, so the gap may already be larger. Epoch AI calls its results “reasonable but approximate measurements based on the best available data,” since the sample is limited.

The headline number needs a qualification. This is the market price of reaching a defined benchmark result, not the cost of doing useful work in the real world. It also does not measure only improvements to algorithms or model architectures.

Earlier capabilitycheaper to repeat
Best availablecostlier per query

Algorithms explain only part of the decline

Hans Gundlach and his MIT colleagues analyzed pricing data from Artificial Analysis between April 2024 and November 2025. Their dataset covered far more models on each test than Epoch AI’s study.

MIT found a 5–10-fold annual decline in prices overall. After accounting for cheaper hardware and competitive pressure, the gain attributable to algorithmic efficiency was about threefold per year.

The difference between the studies is partly explained by reasoning models. They can spend additional computing resources on difficult questions during evaluation, raising the cost of a correct answer even as the price of an individual token falls. Since providers bill for tokens, token prices alone are an incomplete measure.

MIT separated the effects of competition by looking at open models. The remaining change was its estimate of pure algorithmic efficiency. Epoch AI’s figure is 13 times higher because it includes hardware and competition.

More computing can look like better efficiency

MIT also found that some performance gains come simply from allocating more computation to each question. A new model that beats its predecessor on GPQA Diamond may appear more efficient, while actually producing a better score at a higher per-query cost.

The same pattern appears in programming and mathematics benchmarks, though less strongly.

Breakdown of the 2024-2025 price decline into algorithmic efficiency, hardware improvements, and competition. | Image: Gundlach et al.

Breakdown of the 2024-2025 price decline into algorithmic efficiency, hardware improvements, and competition. | Image: Gundlach et al.

Source: the-decoder.com

A benchmark score combines several sources of progress:

Better training
Better data
Improved architecture
More computation during testing

That makes the score useful but blunt. It cannot show which ingredient produced the improvement.

There is also the risk of “benchmaxxing”: optimizing a model for known tests without producing comparable gains on practical tasks. Epoch AI tries to reduce that risk with Mystery Game Puzzles, based on a game whose rules were kept secret. Prices decline most slowly on that test.

That result is consistent with the benchmark-optimization hypothesis, although the format of the task or noise in the data could produce the same pattern. My read is that this is the more important warning in the research: falling benchmark prices are real, but they are not yet a clean proxy for falling costs in deployed systems.

The missing price is the cost of being wrong

Artificial Analysis compares models on quality, price, latency, context-window size and output speed. Those dimensions matter because the cheapest model is rarely the best choice across all of them.

A low-priced model with high latency is a poor fit for a real-time chatbot. A powerful reasoning model may be too slow for an automated workflow. A more expensive frontier model can still reduce total spending if it produces correct answers more often and avoids repeated attempts.

The announcement is quiet about the question buyers ultimately care about: how much does a completed piece of work cost, including errors, retries and waiting time? The benchmark data can show that yesterday’s capability is cheaper to reproduce. It cannot, by itself, show whether a production system has become cheaper to operate.

The economics of tokens were examined in detail in Frontier Radar #3. The broader conclusion is less tidy than the headline: AI capabilities are becoming dramatically cheaper to replicate, while the most capable systems may be charging more computation for each successful answer.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X