Urgent.News

What's breaking now, across thousands of outlets.

AI

Oxford lets OpenAI train its AI models on Bodleian library

University staff voice concerns over reputational risk of partnering with company behind ChatGPT The University of Oxford has allowed the company behind ChatGPT to train its AI models on historical texts from its Bodleian library, as tech companies scour academic institutions for fresh data. The Bodleian material digitised by OpenAI has been used to “populate the OpenAI training set”, according…

Oxford lets OpenAI train its AI models on Bodleian library

The University of Oxford has entered into a partnership with OpenAI, allowing the latter to train its AI models on historical texts from the Bodleian Library. This move comes as tech companies seek fresh data for their AI systems. The Bodleian Library's digitized material has been incorporated into OpenAI's training set, as per internal documents.

The university announced the partnership in March 2025, stating that the digitization process would make the content more accessible to students and researchers. However, the announcement did not explicitly mention that the material would be utilized for training OpenAI's models, which learn by being fed vast amounts of data and recognizing patterns in words.

OpenAI expressed pride in contributing to preserving historical knowledge for the future, emphasizing that the technology is used by over a billion people in everyday life. The partnership's potential impact on the university's environmental commitments and its reputational risk were also noted by university staff, who raised concerns through meeting minutes obtained via a freedom of information request.

Booksellers have reported increased orders for obscure titles, such as a guide to agricultural implements in 18th-century Africa or biographies of 1950s car drivers, speculating that these titles represent fresh data for the next generation of AI models.

The Bodleian Library has retained the rights to the scanned materials, which will be made openly available online in the coming months. The university's primary interest was digitization, but staff acknowledged that the project would also contribute training data to OpenAI. This collaboration is unique, as the university is the only UK member of the NextGenAI project, involving various US research libraries.

By June 2025, 125,000 images from the Bodleian collection had been shared with OpenAI, including 19th and 20th-century PhD theses. Other digitized texts include 10,000 16th-century "broadside ballads" and 18th-century Irish state papers, private letters of a famous Irish novelist, and Dorothy Hodgkin's penicillin notebooks. The university emphasized that the digitization scale is modest and that only out-of-copyright materials were involved.

Written by urgent.news from Guardian Technology's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at theguardian.com →

More in AI

More from Saturday 26 September →