i
DATAIST
News · 2026-09-04

Meta's top Muse Spark 1.3 scores come from a configuration you can't buy

@neuronium_ai @neuronium_ai

Meta released Muse Spark 1.3 yesterday, and the numbers it put at the front of the announcement belong to a configuration nobody outside a partner preview can call. The max setting is still clearing additional safety checks and will arrive "soon." Artificial Analysis, which benchmarked it under limited partner access, currently lists no API provider for it at all. What developers can actually use this week, through the Muse Code shell and the Meta Model API, is xhigh — frontier-class by any reasonable reading, and behind max on most of the charts Meta chose to lead with.

Cover: Meta's top Muse Spark 1.3 scores come from a configuration you can't buy

Meta released Muse Spark 1.3 yesterday, and the numbers it put at the front of the announcement belong to a configuration nobody outside a partner preview can call. The max setting is still clearing additional safety checks and will arrive "soon." Artificial Analysis, which benchmarked it under limited partner access, currently lists no API provider for it at all. What developers can actually use this week, through the Muse Code shell and the Meta Model API, is xhigh — frontier-class by any reasonable reading, and behind max on most of the charts Meta chose to lead with.

Mark Zuckerberg, Meta's co-founder and chief executive, said the model delivers frontier-level performance "almost too cheap to meter" and called it the company's largest update yet for coding and agentic work. On the independent measure, xhigh scores 61 on the Artificial Analysis intelligence index against 62 for max. That puts the shippable version level with GPT-5.6 Sol max, Grok 4.6 high and Claude Opus 5 high. Anthropic still sets the ceiling: Claude Fable 5.1 scores 66 at max and 65 at xhigh, and Claude Opus 5 scores 63 in both configurations.

The gap between Meta's two configurations is uneven, which is what makes the staging matter. On GDPval-AA v2 the company reports 1754 Elo for max against 1709 for xhigh. On OSWorld 2.0 it is 66.9 against 57.2 — nearly ten points. On JobBench, 64.9 against 61.2. Elsewhere the difference collapses or reverses: both configurations score 89.4 on DeepSearchQA, and on Terminal-Bench 2.1 the available xhigh beats max, 89.2 to 88.8.

Meta publishes both sets of figures in a detailed evaluation report, so the deployable version is not hidden. It is simply not the one the launch materials foreground. For an enterprise buyer the useful question is therefore not whether Muse Spark 1.3 reaches the frontier, but how close the version they can deploy today sits to it and what running it costs in practice.

Against its own predecessor the jump is real. Muse Spark 1.2 landed last month; on Terminal-Bench 2.1 it scored 82.9% while Opus 5 scored 86.7%, and across the other headline coding comparisons Meta presented, the older model also stayed behind Opus. VentureBeat's assessment at the time was of a credible coding competitor that still generally trailed Anthropic's best. Muse Spark 1.3 now trades places with OpenAI and Anthropic across coding and agentic benchmarks rather than sitting a tier below them.

Meta also says the base model is easier to steer. It reports training 1.3 to hold several workflows across a long conversation, gather context through tools, find gaps in its own plans, ask the user for clarification when it needs it, and seek confirmation before consequential actions. In internal comparisons by Meta engineers on coding tasks, the model used roughly 20% fewer tool calls and 25% fewer tokens than 1.2. For anyone paying for thousands or millions of agent loops, that kind of behavioural change can matter more than another index point.

Prices did not move. The Standard tier sits where Muse Spark 1.2 left it: $1.25 per million input tokens, $4.25 per million output tokens, $0.15 per million cached input tokens. Meta also kept its unusually cheap Participant tier at $0.10 and $0.20 per million, in exchange for permission to use prompts and responses for training — which, as VentureBeat noted when covering 1.2, is workable for prototyping and a governance problem for anyone with proprietary code or sensitive internal data.

Artificial Analysis measures xhigh at 235.2 output tokens per second, with an average task in its intelligence index costing $0.55. At 61 points that is the lowest task cost of any model it currently measures at that level of intelligence. But Muse Spark 1.2 cost $0.40 per task at 57 points. Per-token prices held flat, and the cost of doing an average unit of independent benchmark work still rose by roughly 38% between generations. Artificial Analysis attributes that mainly to the new model consuming more input tokens on agentic tests.

That does not contradict Meta's 25% figure — Meta is comparing its own coding workflows, Artificial Analysis a broader set of reasoning and agent tasks — but the two together show how little the word "cheap" now carries. Read carefully, Zuckerberg's phrase is not a claim about token prices, which did not change; it is a claim about how much work Meta believes developers will get for the money. The one independent attempt to measure that moved the wrong way. My reading is that per-token pricing has stopped being a useful signal in this category: the bill is set by token rates, reasoning intensity, turn count, tool calls and retries, and a vendor can genuinely improve four of those while the total goes up.

Meta's chief AI officer, Alexander Wang, was blunter about the launch. After Artificial Analysis published its Muse Spark results, Wang posted them on X and added: "Gemini who?" The timing was pointed, because Google shipped Gemini 3.8 Flash the same day, positioned for roughly the same work — long-running software development, autonomous agents and multi-step professional reasoning. Google calls it its best Flash model for reasoning and coding. It is the company's third Flash release in six weeks.

The independent numbers give Wang something to stand on, but nothing like a rout. Gemini 3.8 Flash in boosted reasoning mode scores 59 on the intelligence index at $0.58 per task, against Muse Spark 1.3 xhigh's 61 at $0.55. Two index points and three cents. Google wins clearly on generation speed: Artificial Analysis measures Gemini 3.8 Flash at about 305 output tokens per second against 235 for Muse Spark, roughly 30% more. Google is also cheaper per token right now, at $0.75 in and $3.75 out on a promotional rate that runs to 31 December, against Meta's $1.25 and $4.25 — though the promotion expires into $1.50 and $7.50, above Meta on both sides.

"Gemini who?" is an odd line to aim at a model two points behind you and 30% faster than you. For an enterprise architect the honest version is duller: Gemini is quicker, Muse is slightly stronger as a high-reasoning agent on this one independent measure, and the choice turns on which of those your workload is bounded by.

For some developers the more consequential question has nothing to do with this week's benchmark scoreboard. When Meta launched Muse Code and Muse Spark 1.2 in August, both arrived as closed, API-only products — a visible turn for a company that spent years arguing open models were the main path forward, and that built Llama's reach on downloadable weights. Five days later it turned again. On 10 August Meta released Muse Glimmer, 30 billion parameters, under the Apache 2.0 licence, and Zuckerberg said Meta would open the weights of Muse Spark 1.2 "in the coming weeks." Reuters separately reported the plan to release the 1.2 weights.

Muse Spark 1.3 now arrives as another closed model. That alone does not break the promise; "in the coming weeks" can describe a stretch longer than three weeks. What is harder to account for is the language. Meta's new publication does not mention Muse Spark 1.2 at all. The roadmap instead refers to "an open weights release of Muse Spark," with no version, no date, no parameter count and no licence; Zuckerberg wrote on X that "Muse Spark open weights releases" are coming soon. A specific commitment attached to a version number has become a plural noun attached to none, and nothing in the announcement acknowledges that the earlier commitment was ever made.

Muse Spark 1.3 shows Meta can now iterate closed, frontier-class models on something close to a monthly cadence. The configuration you can actually buy is fast, competitively priced and closer to the top of the independent rankings than anything Meta has shipped before, and the max preview shows how much further the family goes when more compute is spent on reasoning. What Meta still cannot tell a company is which version it will be permitted to run on its own hardware, under what licence, or when. Teams that picked Llama precisely so they could download the weights, deploy on their own servers, tune the model and control their inference economics are being asked to plan around a sentence with no version number in it.