Why old books matter to AI companies
Online collections are becoming harder to rely on for training as they accumulate material generated by AI. Physical libraries offer a different source: texts that may never have been digitised.
Oxford is not OpenAI’s only library partner. The company has similar agreements with the Boston Public Library, Caltech, the Massachusetts Institute of Technology and the University of Michigan, all participants in the NextGenAI project. Oxford is the project’s only British member.
The Bodleian material already shared with OpenAI includes:
Staff also discussed digitising 18th-century Irish state papers, letters by Irish writer Maria Edgeworth and Dorothy Hodgkin’s notebooks on penicillin. Meeting records mention a possible chatbot called Ask the Bod.
What the deal does not settle
Oxford says the digitisation is “modest” and limited to material not protected by copyright. The Bodleian keeps the rights to the scans and plans to publish them online in the coming months. The university also says the project’s use by OpenAI is non-exclusive, and that machine-learning use was discussed openly with staff and students.
That distinction matters beside other ways AI companies have obtained books. Anthropic spent tens of millions of dollars buying books, cutting off their spines to scan them and then sending the volumes for destruction. The company said it did not buy or destroy rare and antique books. 404 Media also placed a tracking device in an ordered used book and followed it to an Amazon facility in the United States, where books were dismantled and scanned.
Oxford’s books stay where they are. But keeping the originals intact does not answer every concern recorded in the university’s meeting minutes: staff raised reputational risks and asked how a deal involving energy-intensive technology would sit with Oxford’s environmental commitments.
OpenAI says it is proud to help modern models preserve historical knowledge. A company representative said more than a billion people use the technology in everyday life, making it important for models to reflect different cultures, histories and perspectives.
I think the sharper test is not whether the scans remain available to the library; Oxford says they will. It is whether open access to the digitised material meaningfully balances the value OpenAI gets from training on it. The announcement leaves that question unresolved, even as the collection becomes easier for the public to consult.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X