Meta has shipped Muse Spark 1.3 through Muse Code and its model API, and the number at the top of the scorecard belongs to a tier customers cannot buy. The max variant scores 62 on Artificial Analysis's Intelligence Index, but it sits in limited preview for partners while it waits on additional safety checks. What users actually get is xhigh, at 61. Token pricing is unchanged from the previous release — $1.25 per million input, $4.25 per million output — and among models scoring 59 or above, nothing runs cheaper.
The cost claim is the strongest thing Meta has here, and it holds. One run of the Intelligence Index suite costs $0.55 on Muse Spark 1.3. Models posting comparable scores charge between $0.94 and $1.23 for the same work. The comparison that flatters it less runs the other way: version 1.2 completed that run for $0.40. Identical token prices and a bill roughly 38 percent higher for the same benchmark means one thing — the model is thinking longer. Efficiency did not improve. The score did, and reasoning tokens paid for it.
The series has moved quickly on the index. The release before 1.2 scored 53, 1.2 scored 57, and 1.3 now posts 61 on xhigh and 62 on max. The composition of that climb is worth reading. GDPval-AA v2 carries 20 percent of the final score, Terminal-Bench 2.1 carries 16 percent and τ³-Bench Banking 14 percent — and those three are precisely where Meta gained the most. Progress shaped exactly like the scoring function is progress that has to be checked test by test.
τ³-Bench Banking puts an agent in a simulated banking environment and makes it use tools. Max scores 52 percent, up from 35 percent for 1.2, which Artificial Analysis currently ranks first in the field. It is also the only test where Muse Spark leads at all. The xhigh tier reaches 47 percent, level with Claude Fable 5.1 (max) and GLM-5.3-Flash and ahead of neither.
Terminal-Bench 2.1 measures coding at the command line. Here xhigh went from 80 to 85 percent and max reaches 86. Claude Fable 5.1 remains well clear at 91.4 percent on max, 91.0 on xhigh and 89.9 on high — its third-ranked configuration still beats Meta's best.
GDPval-AA v2, the heaviest component of the index, moved from 1615 to 1709 on xhigh and 1754 on max. The scale is calibrated against people: 1000 corresponds to what human experts achieve on 220 real professional tasks. Claude Fable 5.1 (max) sits at 1853. Meta's own gap between its two tiers is bought rather than engineered — max spends 62 percent more reasoning tokens than xhigh.
Source: the-decoder.com
On GPQA Diamond, which asks expert-level science questions, Muse Spark climbed from 90 to 94 percent. That is inside the leading group but behind Gemini 3.8 Flash (high) at 95.3 percent and Grok 4.6 (high) at 94.9. CritPt, a research physics test, produced the larger jump and the larger remaining distance: 18 to 26 percent, against 32.3 percent for GPT-5.6 Sol (max) and 31.1 for Claude Fable 5.1 (xhigh).
Two numbers went backwards. AA-LCR fell from 83 to 79 percent. Factual accuracy on AA-Omniscience dropped by as much as three points, because the model now declines to answer more often when it is unsure.
That second regression is the most interesting line in the report, and it is filed as a loss. A model that refuses when uncertain is doing the thing every enterprise buyer says they want, and the benchmark charges it for the privilege. Taken together, the scorecard describes a company that has learned to optimise well: gains concentrated in the three heaviest tests, a first place in the one benchmark where tool use inside a narrow domain can be trained hard, and a tier hierarchy where the top rung is purchased with 62 percent more tokens. That is competent product work. It is not the same thing as closing the distance to Claude Fable 5.1, which still leads on both the heaviest test and the coding test by margins that more thinking time has not touched.
Neither Meta nor Artificial Analysis has named a price for max, the tier carrying the 62. Larger models and an open-weights version are still promised. Being the cheapest route to 61 points is a real position in this market, and it is the position Meta actually occupies today — which makes the number to watch not the index score but the one that eventually appears next to max, a tier whose advantage has to be paid for in tokens on every single call.