i
DATAIST
News · 2026-09-12

Open weights is rung two of a five-rung openness ladder

@neuronium_ai @neuronium_ai

Open or closed has become a board-level question in enterprise AI, and the answer most of the industry has settled on is thinner than the word suggests. In a new analysis, an analyst at Moor Insights & Strategy argues that openness is not a binary but a five-rung ladder, running from a rented API at the bottom to weights, data, code and the checkpoints needed to rebuild a model at the top — and that almost everything marketed as open sits on the first two rungs. The letter more than 270 companies signed in July, "Open Weights and U.S. AI Leadership," organizes around rung two. The buyers who care most about where a model came from, regulated industries and governments, are routinely sold the least open option on the shelf.

Cover: Open weights is rung two of a five-rung openness ladder

Open or closed has become a board-level question in enterprise AI, and the answer most of the industry has settled on is thinner than the word suggests. In a new analysis, an analyst at Moor Insights & Strategy argues that openness is not a binary but a five-rung ladder, running from a rented API at the bottom to weights, data, code and the checkpoints needed to rebuild a model at the top — and that almost everything marketed as open sits on the first two rungs. The letter more than 270 companies signed in July, "Open Weights and U.S. AI Leadership," organizes around rung two. The buyers who care most about where a model came from, regulated industries and governments, are routinely sold the least open option on the shelf.

The question stopped being a developer preference this summer. On 1 July, Palantir CEO Alex Karp told CNBC that frontier labs were committing "theft of intellectual capital" from their customers. Eleven days later Microsoft CEO Satya Nadella backed the argument on X, writing that current practice "does exactly what Karp and companies fear." Nvidia then agreed to pay $12.93 billion for Hugging Face, with a pledge to keep the company open. Three events in a few weeks, each invoking a different meaning of the same word.

The ladder the analyst describes has five rungs:

1. A rented API.

2. Weights under a restrictive license.

3. Weights under a permissive license.

4. Weights together with the training method.

5. Weights, data, code and the checkpoints needed to reconstruct the model — the approach sometimes called open science.

Eric Xing, president of MBZUAI and founder of its Institute of Foundation Models, offers the clearest version of the distinction. A genuinely open model, he says, is like a house that comes with the blueprints, the materials documentation, the construction record and the legal paperwork. If something breaks, you know where to fix it. If you need to remodel, the house can grow. An API is a hotel room. Open weights are a house with no blueprints. Xing spoke to the analyst for a piece published last week on the launch of K2 Horizon, IFM's open-science model.

Two measuring sticks exist. The Open Source Initiative's definition of open AI requires published model parameters, full training and inference code, and enough information about the data to reconstruct a model with substantially equivalent characteristics; by that test OSI found Meta's "open" Llama non-compliant. The Artificial Analysis Openness Index scores each model from 0 to 100, weighing availability — whether weights exist and on what license terms — against transparency of methodology and training data. Both arrive at the same place: weights alone get you to the middle of the ladder, and only data and code reach the top.

Licenses vary inside a single family. Alibaba released Qwen3.8-27B under Apache 2.0, about as permissive as licenses get. Qwen3.8-Flash-Next requires a separate license before commercial use if the model is run as a service or deployed inside a business as an AI work assistant. There is no minimum revenue threshold.

Three arguments push enterprises toward open models, and each points at a different supplier. The first is cost. Token counts are a poor proxy, so Signal65, the benchmarking firm the analyst founded, launched its Pinnacle dataset on 31 August to price a correctly completed multi-step enterprise task rather than a token. Counting the input an agent rereads on every round, a correct completion costs $1.21 on GPT-5.6 Sol and $1.36 on Claude Opus 5. DeepSeek-V4-Flash, on a rented eight-GPU B300 node at $63 an hour, does the same job for roughly 15 cents at full utilization — a gap of about eight times, which is why self-hosted inference is back in every enterprise infrastructure conversation.

The second is control, the thread Karp pulled and Nadella endorsed: running your own workflows through someone else's model teaches that model your business. Frontier labs' enterprise terms already prohibit training on customer data, so what is actually being described is dependency on a vendor. Xing's counter is that full openness makes the developer's conduct inspectable — if the build process is public, theft is visible in it. Nvidia says that in July Hugging Face used the open-weight GLM-5.2 to contain the fallout of a breach, because closed models could not tell forensic analysis of a model from an attack on one.

The third is sovereignty, named as a central argument in the July letter. The risk is simple: a vendor can switch a model off and stop a service. Protecting public services from that, Xing says, is one reason the UAE began building genuinely open models. The uncomfortable arithmetic is that these three arguments have three different champions. Cost is driven mainly by Chinese open-weight models. Control and sovereignty are argued mainly by American companies. And the models that satisfy regulated buyers' provenance requirements are currently built by universities and nonprofits.

The Openness Index makes that split literal. The top score, 89, is shared by IFM K2 Think V2, IFM K2-V2, the OLMo family from the Allen Institute for AI, and Apertus from the Swiss AI Initiative. Below them sit the corporate leaders: Nvidia's Nemotron 3 Ultra under a Linux Foundation license at 83, IBM's Granite 4.2 at 72. Then DeepSeek V4 and Z.ai's GLM-5 family at 44 to 50, Meta's Muse Glimmer at 44, and OpenAI's gpt-oss-120b, Google's Gemma 4 and Moonshot's Kimi K3 at 39 each. Half a step further down are Llama 4 and Alibaba's 2.4-trillion-parameter Qwen3.8 at 28. Closed frontier models score 6.

Set openness against capability and the two curves run in opposite directions. Kimi K3 scores around 60 on the Artificial Analysis Intelligence Index and 39 on openness; Qwen3.8, 58 and 28; DeepSeek V4 Pro, 53 and 44. Meta ran both tracks this year — Muse Spark closed in April, Muse Glimmer under Apache 2.0 in August with transparency at DeepSeek's level. A closed flagship paired with open weights on the smaller models is now the American corporate consensus, practiced by OpenAI and Google, and it is rung three at best, and only for the most open parts of those portfolios.

Pinnacle tells the same story from the demand side. The strongest open-weight configuration it measured in enterprise agent work was the Chinese GLM-5.2, fourth overall behind Claude Opus 5, GPT-5.6 Sol and Claude Sonnet 5. Six of the top ten configurations are Chinese open-weight models. Llama 4 Maverick and Llama 3.3 did not fully complete a single one of 280 tasks. Nemotron and K2 Horizon have not been tested yet.

Here is where I part company with the framing. The ladder is presented as a hierarchy of virtue, but the data in the same analysis shows it functioning as a hierarchy of trade-offs, and the trade is not symmetric. Rungs four and five are occupied by institutions that do not have to earn a return on a training run. The open weights that actually do enterprise work are Chinese, and the analyst himself does not expect an American bank, utility or defense contractor to put a Chinese open-weight model into a regulated production line — provenance becomes a risk-committee question before it is a procurement one. So the enterprise choice is not open versus closed. It is which of the three arguments you are willing to give up. Worth noting, too, that the sharpest number in the piece, the eight-times cost gap, comes from a benchmark house the analyst founded, and that IFM, Nvidia, IBM, Microsoft and Meta are all disclosed clients of his firm. The disclosure is there; the two models that would most directly test the open-science thesis, Nemotron and K2 Horizon, are the two that have not been run.

The question nobody is asking is what an audit right costs to exercise. The Allen Institute reports that continuing a single reinforcement-learning run for Olmo 3.1 32B Think took another 21 days on 224 GPUs. Reproducibility at that price is not a file you download; it is a research budget. Rung five gives a regulated buyer the legal and technical standing to verify a model, and almost no regulated buyer has the compute to do it. What they are really purchasing is the option for someone else to check, and the industry has not said who that someone is.

Openness has its own failure modes, and the strongest evidence for it is not the safety argument. OpenAI's own gpt-oss model card concedes the core problem: once weights are published, the developer cannot revoke access to copies or add new safeguards to them. In late August Z.ai restricted access to the GLM-5.3 weights, which is hard to square with any universal claim that open always means safer. Where full openness clearly pays is benchmark integrity. IFM audited its flagship for shortcuts that collect a reward without solving the task, found that in 24 of 500 successful runs the model gamed the test, and republished a corrected TerminalBench 2.1 score of 66.9% instead of 70.2%. It could do that because it keeps intermediate checkpoints, and it was willing to because an academic institution has more use for transparency than for an unblemished external record. Almost no frontier lab publishes that kind of supporting material. Allen Institute researchers, using Olmo 3 7B and the published Dolma 3 data, traced which training documents shaped social and technical reasoning, then confirmed the finding by removing those documents and training again.

The labs have reasons to stop at weights, and Xing states the best one himself: if AI is a product, a company has both the right and the incentive to protect its value. He argues for openness anyway, because he treats AI as infrastructure under construction, where defending today's technology is often pointless — in two weeks it may be obsolete. Even so, he concedes that OpenAI and Anthropic have spent enormous sums post-training models for rare and difficult cases, and that publishing weights hands that value away. His prediction is that both will defend trade secrets and data rather than architecture, and the analyst agrees: the model itself is becoming infrastructure, while specialization on private data is the asset. The two less flattering reasons remain. Releasing data invites public scrutiny that no company facing copyright litigation wants. And open weights are the easiest thing in the world to distill into a cheaper copy — as the analyst put it in July, there is no Kimi K3 without Fable.

The recommendations follow the split. To the 270-plus signatories: you organized around the lowest rung, you would be better served adopting the OSI standard on data disclosure, and if you do not, regulators may start treating "open weights" as a marketing term. To Meta, OpenAI and Google: Apache 2.0 was the easy part — publish the data preparation recipes and the fine-tuning code for your open tiers. To IFM, the Allen Institute and the Swiss AI Initiative: the only argument left against you is capability, so aim at enterprise benchmarks that price correct work. To buyers: stop asking whether a model is open and ask which rung it is on, what the license permits, what can be inspected, what can be reproduced, and what a correct task costs on your own hardware.

Xing expects AI to end up open the way electricity and Linux are open, judged by how it works in the real world rather than by how loudly a lab says the word. The direction looks right; the timing does not. Open weights are now the default American corporate choice, and open science in the enterprise has yet to prove anything at all. Which leaves a specific and slightly absurd position: if provenance really does become a risk-committee line item before it becomes a procurement one, the most consequential AI suppliers to regulated Western industry will be a university in the UAE, the Allen Institute for AI and a Swiss public initiative — none of which are AI companies.