Artificial Analysis has rebuilt its Intelligence Index after the score it gave OpenAI's GPT-6 Astra was met with skepticism. Several other evaluations, OpenAI's own tests among them, had shown Astra clearly ahead of rival models; Artificial Analysis rated it level with its own predecessor. Under the revised index Astra gains four points on that predecessor and lands second, behind Anthropic's Claude Fable 5.1 and ahead of a model from Meta.
The methodology changed in four ways. Two benchmarks were added: AA-Briefcase, which tests the knowledge and skills needed for real work tasks, and GDP.pdf from Surge AI, which measures analysis of PDF documents. GPQA-Diamond was dropped, because models now solve it and it no longer separates them. The share of the final score coming from private test data rose to 40%, which is meant to make results harder to game. And Artificial Analysis fixed counting errors in several benchmarks and revised its grading systems for more stable scores.
The organization's account of the timing is that it had been deliberately holding updates back to keep scores comparable through major model launches, and that the top of the leaderboard started moving fast enough to require an interim release. Version 5 has been in development for eight months and will ship in stages.
Two other results arrived with the revision. Artificial Analysis says Astra spends fewer tokens per task than any other frontier model. On cost against performance, first place is shared four ways, by Anthropic, OpenAI, Meta and Zhipu AI.
The awkward part is the sequence. A benchmark that changes its methodology after the scored party and much of the field disputed a specific result is in a position where the correction and the capitulation look identical from outside, and Artificial Analysis does not have a way to prove which one this was. The eight months of development on version 5 are the strongest evidence for the innocent reading: fixes that take that long were queued before the Astra argument started. What the argument appears to have changed is the release date, not the content — and that is a smaller sin than it first looks, since an index that knowingly publishes numbers it believes are wrong in order to protect comparability is making a worse trade than one that ships the fix early.
The more useful defense is in the result itself. If the revision had been shaped to end the complaints, Astra would be first. It is not. Four points and second place, behind a competitor's model, is not what a captured benchmark produces.
What is missing is the breakdown. Astra gained four points across a scoring system that simultaneously added two benchmarks, removed one, reweighted toward private data and corrected arithmetic errors — and Artificial Analysis has not said how much of the gain comes from which change. That matters because the individual edits pull in different directions. Removing a saturated test mechanically widens the spread at the top for everyone, not just for Astra, so part of that four-point move may be an artifact of the field decompressing rather than Astra being re-measured. Adding a document-analysis benchmark rewards whatever models happen to be good at documents. Without per-component numbers, the four points are a headline, not a finding.
The change that deserves the most attention got the least. Pushing private data to 40% of the final score is the only edit here aimed at the structural problem, which is that public benchmarks leak into training sets and the leaderboard slowly stops measuring capability and starts measuring exposure. Removing GPQA-Diamond because models have solved it is the same problem arriving at its terminal stage, described as a routine housekeeping item.
Independent indices are the only public referee the model market has, since every lab grades its own homework and publishes the score. This one has now revised its rules mid-season after the largest player it scores said the rules produced the wrong answer, and it did so without publishing the component-level arithmetic that would let anyone check the revision on its merits. Astra ending up second rather than first is the reason to extend Artificial Analysis the benefit of the doubt. It is not the same thing as being able to verify it.