i
DATAIST
News · 2026-09-18

Microsoft's own 93% figure lands in the New York Times case

@neuronium_ai @neuronium_ai

Microsoft's own director of applied research, Brent Hecht, called the collection of published work for AI training "the largest theft of labor in human history" and a "staggering theft of unprecedented scale." He wrote it in an internal memo in January 2023, and it became public only now, in unsealed material from the copyright suit The New York Times brought against OpenAI and Microsoft three years ago. The same filings show OpenAI's head of ChatGPT describing his own products as an "existential threat" to publishers, and Microsoft's own measurements showing that Copilot cut referral traffic to the Times domain by as much as 93% compared with traditional Bing search.

Cover: Microsoft's own 93% figure lands in the New York Times case

Microsoft's own director of applied research, Brent Hecht, called the collection of published work for AI training "the largest theft of labor in human history" and a "staggering theft of unprecedented scale." He wrote it in an internal memo in January 2023, and it became public only now, in unsealed material from the copyright suit The New York Times brought against OpenAI and Microsoft three years ago. The same filings show OpenAI's head of ChatGPT describing his own products as an "existential threat" to publishers, and Microsoft's own measurements showing that Copilot cut referral traffic to the Times domain by as much as 93% compared with traditional Bing search.

Start with what these documents are, because it matters. Most of the new material sits in a memorandum written by the Times, not in the exhibits themselves, which remain sealed. The quotations in it are therefore published without their original context, selected by the side that benefits from them. That is not a reason to dismiss them. It is a reason to read the internal opinions as opinions and the internal numbers as numbers.

The numbers are the part that should worry the defendants.

Copying was never seriously the contested question in this case. Fair use was. The legal status of training a generative model on copyrighted work is still unsettled, and judges have so far leaned toward the AI companies' argument that training is transformative. Earlier this month the Trump administration filed in support of OpenAI's position that copyrighted material can be used to train large language models without a license. On the transformation question, OpenAI and Microsoft have been winning.

But fair use also turns on whether the new use substitutes for the original and damages its market, and that is where Microsoft's internal record now works against Microsoft. The 93% figure is Microsoft's own. In a January 2024 internal presentation, Hecht called the decline a "vicious circle" that degrades both Microsoft's models and the internet they are trained on. One Microsoft document states plainly that the company's end product is capable of threatening the economic base of the suppliers it depends on — in the large language model business, the content supply chain. Another warns of a real risk that generative AI substantially disrupts the employment of the people who produced the data the base model learned from.

Then there is the substitution argument, which the defendants' own executives make better than the plaintiffs could. Nick Turley, who runs ChatGPT at OpenAI, wrote internally that products like the chatbot pose an existential threat to publishers because they substantially replace the original material, and will replace it more often as they improve. OpenAI president Greg Brockman described the models as working particularly well with news. Satya Nadella, deposed earlier this year, agreed that a conversation with a chatbot substitutes for the click to the source site, because the user gets the information on the AI platform instead.

Nadella went further. He testified that anything behind a paywall should be licensed by whoever wants it for context or for training, and that had he known OpenAI was collecting paywalled data and training on it, he would have used Microsoft's right to require OpenAI to retrain those models.

On scale, the filings supply the first hard counts. OpenAI's intermediate training sets alone contained more than 91,692 copies of works published by the Times, the Daily News and the Center for Investigative Reporting. A dataset built from Common Crawl included more than 2 million documents from nytimes.com alone. Data gathered under Project Mango — one of two pipelines, alongside Project Taxi, through which Microsoft passed training data to OpenAI — was compiled, the plaintiffs say, into a training set holding copies of at least 160,903 unique works from news publishers. OpenAI handed Microsoft the complete GPT-3 training set, which Microsoft used to assess how to put OpenAI models into its own commercial products.

The collection methods are described in the same detail. Content was pulled from the Bing index. OpenAI staff allegedly built a way to get past paywalls without being noticed; when researcher Nick Ryder told Brockman he had found a way around the nytimes.com paywall, Brockman replied that this was good. Staff allegedly assembled the WebText and WebText2 training sets with a disproportionate share of scraped news, and pulled millions of articles out of Common Crawl. And the filings describe deliberate stripping of copyright notices from training data before it reached the model, because the researchers did not want the model reproducing those notices for users.

Here is my read. The quotations are the headline and the weakest evidence — Hecht was an internal critic writing in a private channel, not a policymaker, and courts do not decide fair use by canvassing employee sentiment. The 93% is different in kind. It is not an opinion about harm; it is Microsoft measuring harm, in its own product, with its own instrumentation, and writing the result down. Nadella's deposition is the other tell: a defendant carefully drawing a line between what his company did and what its partner did, under oath, while the joint defense is still notionally joint. Steven Lieberman, counsel for the New York Daily News, said the material shows for the first time that OpenAI and Microsoft understood they were acting wrongly. That is an advocate's framing of documents an advocate selected — but the underlying artifacts are not advocacy, and the defendants produced them.

The question nobody is asking is what happened between January 2023, when Hecht wrote the theft memo, and January 2024, when he was presenting the traffic collapse as a vicious circle. A year separates a warning from a measurement of the thing the warning described, and in that year Microsoft shipped the product that produced the measurement. Nothing in the unsealed material suggests the first document changed the second. That sequence — diagnose, ship anyway, measure the damage — is the part that will read badly in front of a jury, and it has nothing to do with copyright law.

OpenAI and Microsoft did not respond to requests for comment.

If the courts eventually accept the transformative-use argument and reject the market-harm one, the industry gets its legal answer and keeps Hecht's problem: models trained on a content supply chain that the models themselves are draining. The lawsuit can be won without the vicious circle being solved.