i
DATAIST
News · 2026-09-15

Nvidia, Palantir and Booz Allen curb use of Anthropic's Fable

@neuronium_ai @neuronium_ai

Nvidia, Booz Allen Hamilton and Palantir have each restricted how Anthropic's Fable can be used inside their businesses, and the stated reason is the same in all three cases: none of them is confident about what the model does with what it is shown. Nvidia now confines Fable to lower-sensitivity work such as open-source projects and runs its own Nemotron models for internal processes, including AI-assisted supply-chain monitoring. Booz Allen Hamilton has barred employees from using Fable with its own cybersecurity software. Palantir has blocked Fable from launching through the software it ships to its customers.

Cover: Nvidia, Palantir and Booz Allen curb use of Anthropic's Fable

Nvidia, Booz Allen Hamilton and Palantir have each restricted how Anthropic's Fable can be used inside their businesses, and the stated reason is the same in all three cases: none of them is confident about what the model does with what it is shown. Nvidia now confines Fable to lower-sensitivity work such as open-source projects and runs its own Nemotron models for internal processes, including AI-assisted supply-chain monitoring. Booz Allen Hamilton has barred employees from using Fable with its own cybersecurity software. Palantir has blocked Fable from launching through the software it ships to its customers.

Justin Boitano, Nvidia's vice president of enterprise AI, said the company believes zero data retention — ZDR, meaning no storage of customer data — should be switched on by default. Nvidia is not a hostile party here. It has invested in Anthropic, plans to keep investing, and supplies the hardware Anthropic uses to develop models. It still will not run its own supply chain through the product.

Bill Vass, Booz Allen's chief technology officer, said the restriction came from a concern that the model could train on the firm's code. Booz Allen was one of the first users of Anthropic's Mythos model.

Palantir has set the hardest condition: no Fable through its client software until Anthropic provides irrevocable zero-data-retention guarantees. At a customer meeting, chief executive Alex Karp said companies are tired of being "exploited" by AI labs. Karp has voiced that distrust before, and it happens to align with Palantir's commercial interest — the company would rather customers reach models through its own platform, which it describes as secure, than go to model vendors directly. That does not make the complaint wrong. It does mean Palantir's security posture should not be read as disinterested.

The labs have moved. In August, OpenAI began letting GPT-5.6 Cyber users keep security logs on their own servers. After customer dissatisfaction, Anthropic followed the same route, rolling out an equivalent program for selected clients this autumn.

What retention terms do not cover is how much the labs learn anyway. According to The Information, OpenAI and Anthropic collect metadata and technical usage information from enterprise customers. OpenAI describes that data as de-identified: stripped of anything that would connect it to a specific customer. Its site says business data is passed through automated classifiers and safety tooling so the company can better understand how its services are used, and that the resulting classifications are metadata about business data rather than the business data itself. Some customers say they are not sure what that metadata contains and that the current level of transparency is not enough.

The reason this stopped being an abstract contract argument is Tristan Buckmaster. The mathematician and his co-author Levent Alpoge used AI models to make progress on the Navier-Stokes equations, uploading their drafts through Codex. Shortly afterwards, OpenAI published a result of its own using the same unusual solution path. OpenAI first said it could not rule out that de-identified data gathered from use of its products had helped improve its models. After an internal investigation it updated the publication to say that Buckmaster's Codex prompts in the two months before the September 8, 2026 publication could not have influenced the system by any route, including training.

John Schulman — an OpenAI co-founder who spent a short period at Anthropic and now works at Thinking Machines — responded by laying out how training on customer data actually happens. At one end is direct pretraining on user data, which carries a high risk of reproducing the source content. Then comes distillation of large models into smaller ones, and the construction of reinforcement-learning tasks from "user traces". His examples of the last category: using explicit user feedback to train a reward model, and ingesting a user's coding environment and commit history to turn them into reinforcement-learning environments. That category has little to do with verbatim reproduction and can still expose a customer's intellectual property. De-identification is unreliable, Schulman said: a user can sometimes be recovered from a small number of bits of data, and it offers no protection against IP leakage at all.

Sara Hooker, an AI researcher previously at Cohere and Google DeepMind, described a related gap. There are sophisticated synthetic-data methods, she said, that preserve confidentiality while reproducing the statistical properties of the original set. A lab need never use the original data directly to extract its regularities and arrive at the same practical result. Hooker's warning to companies sitting on proprietary data was blunt: they have a limited window to build their own intelligent systems on it, or they will be helping a frontier lab that sooner or later competes with them in their own industry.

Schulman later softened his own framing. Training on user data, he wrote, is extremely unlikely to noticeably affect the growth of frontier model capability; scaling pretraining and reinforcement learning do that work. User data is more useful for finding bugs and situations that are hard to reproduce with paid annotators. He added that model developers differ in how aggressively they use customer data, and that repository ingestion is not hypothetical — the obvious reference is coding tools such as Codex and Cursor, where users connect their repositories directly to the service. He called for stricter norms on disclosing how companies train models on user data.

The softening is narrower than it looks, and the distinction matters. Schulman's claim is about capability: customer data will not make the next frontier model meaningfully smarter. That is not what Nvidia, Booz Allen and Palantir are worried about. They are worried that their code, their supply chain and their security tooling end up shaping a system that a competitor — or the lab itself — later runs against them. A statement about capability gains is not an answer to a claim about intellectual property exposure, and the two have been drifting together into a single reassurance.

Which is why the concession the labs made is the wrong shape. Zero data retention is a promise about storage, and storage is the most auditable part of the pipeline and the least of what the customer fears. Everything Schulman described — reward models trained on feedback, reinforcement-learning environments built from commit histories — and everything Hooker described sits upstream of retention or beside it. A customer holding a signed ZDR term has bought a guarantee that the lab will not keep the file. It has not bought a guarantee that the lab learned nothing from it, and there is no mechanism by which it could check either way.

The part of the Buckmaster episode nobody has answered is how OpenAI reached its conclusion, or how anyone outside the company could verify it. The updated statement is a categorical negative — not by any route, including training — about an internal pipeline, established by an internal investigation. It may well be accurate. It is also precisely the kind of claim a customer with no visibility into that pipeline has no way to test, which is the condition Palantir is refusing to accept and the one Schulman is asking the industry to write rules about.

Hooker's clock is the harder problem. If the window she describes is real, the companies with the most valuable proprietary data face a decision no data-handling clause can settle: use the frontier labs now and accept assurances they cannot verify, or spend years building something weaker in-house while the gap widens. Nvidia can afford Nemotron. Most of the labs' customers cannot afford the equivalent, and their contracts are still being written as though the only question were where the logs are kept.