i
DATAIST
News · 2026-09-17

Baseten pitches open-weight safety as 6,000 abliterated models pile up

@neuronium_ai @neuronium_ai

Base Labs, the research group Baseten stood up earlier this year, has partnered with Hugging Face and Goodfire to build safety into open-weight models, and has invited developers from the wider community to help construct the framework. What the three will actually do together is not public: none of them has disclosed the technical scheme. The backdrop is abliteration, a technique for stripping guardrails out of open model weights that has gone from fringe to routine. Hugging Face's own catalog now lists more than 6,000 models that have been through it.

Cover: Baseten pitches open-weight safety as 6,000 abliterated models pile up

Base Labs, the research group Baseten stood up earlier this year, has partnered with Hugging Face and Goodfire to build safety into open-weight models, and has invited developers from the wider community to help construct the framework. What the three will actually do together is not public: none of them has disclosed the technical scheme. The backdrop is abliteration, a technique for stripping guardrails out of open model weights that has gone from fringe to routine. Hugging Face's own catalog now lists more than 6,000 models that have been through it.

Base Labs' stated job is to develop and publish methods for training and monitoring open models. Baseten wants whatever comes out of that to become the standard for open models — transparent, and built into the training and deployment process rather than added on top afterwards. On X the company argued that openness can serve AI safety, because it offers more visibility into how models behave, and more ways to turn safety research into practical and transparent controls, than a closed-source approach does.

Goodfire, replying to that post, stated the goal slightly differently: safety should be built into open models and provided by whoever serves them to users. Goodfire studies the AI black box and explains how models arrive at decisions, which makes it the likely owner of the embedding work.

The money around the partnership is lopsided in a way that describes the industry. Baseten sells inference. It raised $1.5 billion in a Series F in June, taking its valuation to $13 billion. Goodfire raised $150 million in a Series B earlier this year, led by B Capital, to build out its interpretability platform. The company that serves models raised ten times in one round what the company that explains them has raised in total.

Put the two framings next to the problem and they do not quite line up. Abliteration works by removing behavior that already sits in the weights. A framework whose pitch is that safety will be trained in rather than bolted on is proposing to push that behavior deeper into precisely the place the technique reaches. That may well be tractable, and Goodfire's interpretability work is the most plausible route to it — but it is a technical claim, and the technical scheme is the one thing the three companies have not published. Until they do, this reads as a statement of intent with three logos on it.

The harder question is the one Goodfire's own formulation raises. If safety is the responsibility of whoever serves the model to users, the open-weights premise punches a hole straight through it: anyone can download the weights and serve them themselves, which is what open weights means. Enforcement at the serving layer reaches the users who were going to behave anyway and misses everyone who is not asking permission — including whoever pulled down any of the 6,000 abliterated checkpoints already sitting in the catalog. It also puts the control point inside a $13 billion inference provider. I do not read that as cynicism; the serving layer genuinely is where controls can be applied. But it is a commercial position as much as a safety one, and both readings are correct at once.

There is also something worth noticing in the guest list. Hugging Face is the registry hosting the 6,000 stripped models and is now a partner in the effort to make stripping them harder. That is the distribution channel volunteering to be part of the fix, which is more than most distribution channels do — and also the clearest signal of how far the problem has already travelled past the point where a training-time answer helps. Whatever the three build, it arrives after the 6,000, and by construction it can only shape the weights that have not been released yet.