i
DATAIST
News · 2026-09-15

OpenAI Foundation funds a hunt for dead biotechs' drug files

@neuronium_ai @neuronium_ai

OpenAI Foundation, the nonprofit parent of OpenAI, has started paying for biological data to be created. Under a new program called Data for Public Health, its first round awards $40 million to a University of North Carolina at Chapel Hill program gathering data on new cancer vaccines, $500,000 to the advocacy group 1Day Sooner to buy datasets from bankrupt biotech companies, and support to OpenAdmet, which runs competitions on predicting how drugs behave. The premise stated by the foundation is that AI will not deliver major breakthroughs in treating disease until researchers give models far more information than they currently have.

Cover: OpenAI Foundation funds a hunt for dead biotechs' drug files

OpenAI Foundation, the nonprofit parent of OpenAI, has started paying for biological data to be created. Under a new program called Data for Public Health, its first round awards $40 million to a University of North Carolina at Chapel Hill program gathering data on new cancer vaccines, $500,000 to the advocacy group 1Day Sooner to buy datasets from bankrupt biotech companies, and support to OpenAdmet, which runs competitions on predicting how drugs behave. The premise stated by the foundation is that AI will not deliver major breakthroughs in treating disease until researchers give models far more information than they currently have.

The bankruptcy grant traces back to a proposal made last year by Ruxandra Teslo, a clinical trials policy analyst. Her idea was to buy the files of failed biotechs through bankruptcy proceedings: detailed regulatory documents, manufacturing strategies, safety data — material normally treated as trade secrets. She called it the lost archive of biotech and suggested it could train AI able to assist in the opaque process of drug approval. 1Day Sooner, which advocates for clinical trial participants, will do the buying. Teslo advises the group.

The bottleneck claim has backing inside the field. Morgan Levin, former vice president for computation at the longevity company Altos Labs, says data has become the main constraint on applying AI successfully in biology. The foundation's own framing is that many future advances in preventing and treating disease will come from combining new models' capabilities with more observation of the real world.

The $500,000 is a proof-of-concept sum, and Josh Morrison, 1Day Sooner's president and co-founder, describes it as one: the grant is meant to demonstrate that the organization can actually obtain these datasets. He estimates non-exclusive copies of corporate datasets can be bought for a few tens of thousands of dollars each. The group currently holds three. Two were donated by Lumen Bioscience, a biotech that had itself used Chapter 11 to acquire another company's drug development information. Two further attempts to buy pharmaceutical documents this year failed — 1Day Sooner's bids were not accepted.

The files in question are called common technical documents: a company's correspondence with regulators, detailed scientific and medical measurements, effectively everything known about a particular drug. Teslo, who writes for Works In Progress and is a visiting fellow at the Washington think tank Institute for Progress, argues that accumulating them could turn an AI into an expert on drug regulation, which she sees as one of the main ways to get new medicines to market faster. Her point is that the picture of AI being built and then curing cancer sits badly against how drug development actually works: about 70% of the cost and time goes to clinical development — running the trials, verifying the compound — and that process stays almost entirely opaque, particularly for the small biotechs that originate new drugs.

The foundation has money for vastly more than this. It owns 26% of OpenAI, a stake that could reach $250 billion if the company completes an IPO at the $1 trillion valuation now discussed; for comparison, the Gates Foundation and its associated trust held roughly $180 billion in assets at the end of 2025. Against that, the operation is small and new. It is based in San Francisco, still hiring for many key positions, and only began scaling up grantmaking this year. Its largest single gift so far is $100 million, made in August to the Common Health Coalition, which helps patients obtain hepatitis C drugs. Jacob Trefethen, who runs the foundation, says it operates separately from OpenAI but shares the company's official mission, and expects to commit $1 billion in grants by the end of the year.

Locating biology's bottleneck in the data rather than the models is a convenient diagnosis for an organization whose parent sells models. It says the capability is already adequate and that what is missing lies outside the lab. Levin's assessment and the 70% figure on clinical development both suggest the diagnosis is largely correct. It is still worth noticing what follows from it: nonprofit money spent assembling biological corpora is also a subsidy to the input supply of the industry whose shares the foundation holds. Nothing in the program is structured to keep those corpora away from the frontier labs, and nothing in the announcement says who gets to use what the grants produce.

The sharper omission is the one the foundation's own field has spent the past month arguing about. People inside AI companies put the odds of human extinction within the next decade at 10% or higher, with engineered biology among the routes they name. Last week Sam Altman and xAI founder Elon Musk endorsed Anthropic chief executive Dario Amodei's call to slow improvements in model capability so that risk-prevention systems can keep pace. A program whose explicit purpose is to assemble far more high-quality biological data for models to learn from sits at an awkward angle to that position, and the announcement does not acknowledge the tension at all.

There is a consent question underneath as well. Common technical documents hold detailed medical measurements — effectively everything learned about a drug, which includes what was learned from the people it was tested on. Bankruptcy courts are already being described as a new land grab for AI training data: last month Google won the auction for bankrupt airline Spirit Airlines' corporate data, including 100 million emails, over objections from flight attendants and other employees worried about personal or protected information being exposed. A failed biotech's archive raises the same objection with patients standing where the employees stood, and the buyer here is an organization that exists to represent trial participants.

The $500,000 is the more interesting number, not the $40 million. UNC's cancer vaccine data will be produced to order; the bankruptcy archive is a bet that the most valuable biological data already exists and is simply locked inside companies that failed. If Morrison is right that copies go for a few tens of thousands of dollars each, the whole lost archive costs less than a rounding error against a stake that could be worth $250 billion — which suggests the obstacle was never price. Two of 1Day Sooner's bids were rejected this year. Someone on the other side of those auctions has a different view of what the files are worth.