OpenAI's GPT-6 Astra took first place on ErdosBench, ulam.ai's benchmark of 226 open mathematical problems modelled on well-known Erdős problems. It scored 3.23, solved 106 of them, of which 43 were solved completely, and disproved a further 27. The result that matters more than the score, though, is what OpenAI says about how it got there. In an essay titled "Alien Intelligence", chief scientist Jakub Pachocki writes that the company believes it could make its models meaningfully stronger at mathematical research by putting more attention into it — and has decided not to, because work on recursive self-improvement and automated AI safety research is more urgent. The strongest mathematical model currently available is, on its maker's own account, a by-product.
Against its predecessor the gain is real but modest. GPT-5.6 Sol, running in maximum reasoning mode, solved 78 problems for a score of 3.12. Astra writes better scientific prose and makes fewer inflated claims than Sol did; in some cases it understated its own results instead. Przemek Chojecki, who built the benchmark, put the improvement across the various research skills he tested at 5 to 10 percent, and added that ErdosBench is still far from saturated.
GPT-6 Astra leads ErdosBench with a score of 3.23 and 106 problems solved, followed by GPT-5.6 Sol with 3.12 and 78 solved
Source: the-decoder.com
That gap between the headline — first place on a hard open-problem benchmark — and the actual delta is the story mathematicians should be reading. A 5 to 10 percent improvement on an unsaturated benchmark is a breathing spell, not a reprieve.
The deliberate part is what makes Pachocki's essay unusual. Mathematical achievement was front and centre in OpenAI's own first announcement of the model, which makes it strange to then read the chief scientist explaining that mathematics is explicitly not a priority. The resources are going to recursive self-improvement and to protecting future AI systems instead; according to Pachocki, that is the only way to stay at the frontier of AI research. Reports appeared after the essay was published that OpenAI has trained stronger mathematical models internally, but the situation remains complicated, and the RSI and safety agendas have expanded sharply in the meantime.
Read as a disclosure rather than as marketing, the statement says something more consequential than anything in the benchmark table: at the frontier, capability is now allocated. Directions compete for it. A model that dominates in mathematics will not automatically be the best at everything else, because the mathematics came out of a budget that something else did not get.
That framing lines up with a visualisation by Cambridge researcher Adam Hunt, which sets two possible trajectories side by side. The mainstream AGI hypothesis has models improving gradually and broadly until they can handle everything a human can. The alternative is an increasingly spiky path: a model acquires exceptional ability in a few domains — programming and mathematics, say — while stalling or even regressing in language quality, common sense and social reasoning. What comes out the other end is a narrow specialist rather than anything most people would call general intelligence. AGI, of course, can be declared to mean whatever anyone wants it to mean.
Two paths for AI: gradual, broad improvement (top) versus extreme specialisation with stagnating baseline capabilities (bottom)
Source: the-decoder.com
My reading is that OpenAI has just supplied the strongest available evidence for the spiky path, and did so by accident. Not because Astra is narrow, but because the leading lab in the field has conceded in writing that it cannot push everything at once and has to pick. Capability grows where the optimisation goes, and more mathematics means less of something else. The bet on recursive self-improvement is precisely a bet on escaping that constraint — the hope being that a future model performs the optimisation itself and widens its own abilities faster than a research organisation can allocate them by hand, mathematics included. Until that works, the trade-off stands, and every frontier release should be read as a statement of priorities rather than a measure of general progress.
Which leaves mathematics with a problem that does not depend on whether Astra is 5 percent better than Sol or 50 percent better. The field is running into computational capability that helps solve problems without necessarily helping anyone understand them. Terence Tao raised this at the 2026 International Congress of Mathematicians: if models keep producing proofs faster than humans can verify them, mathematics shifts from a shortage of proofs to a surplus of them. The hard task then stops being solving problems and becomes deciding which results actually matter. Tao argues the field faces a crisis of its own values and practices comparable to the upheaval in the foundations of science in the early twentieth century.
For now the hardest problems remain unsolved, and that buys some time. The direction of travel is already visible, though some mathematicians do not believe language models can produce genuine breakthroughs at all, on the grounds that they lack human creativity — the same objection that sustains doubt about whether AI can truly self-improve. If that objection is wrong, the surplus Tao describes arrives on OpenAI's schedule rather than mathematics'. If it is right, then the RSI bet that cost mathematics its optimisation budget was never going to pay out either.