AI companies are acquiring vintage books to ensure clean training data.
The message from data broker ISBNdb is straightforward: the best AI training data available is currently sitting idle on shelves. The downside is that accessing this data requires cutting the spines off millions of books, and the identity of the paying labs must remain undisclosed.
AI is facing a self-created pollution issue. Much of the web now consists of machine-generated content, leading to models potentially consuming their own byproducts. One company believes the solution lies in old printed books released before the chatbot era began.
404 Media reported that ISBNdb is now acquiring physical books in bulk for AI labs to utilize as training data. The company claims to be the largest book database globally. Their website asserts that “the world’s best AI training data is sitting on a shelf,” emphasizing that books are “dense, edited, and authoritative.”
The significance is in the publication date. ISBNdb asserts that books printed before 2022 were produced before the explosion of large language models, indicating they lack AI-generated text. This distinction is crucial due to model collapse, which refers to the decline seen when models are trained on the outputs of previous models, resulting in each generation being slightly inferior.
The open web cannot provide such assurance. A rapidly increasing portion of online content is machine-generated, part of the same mess that AI labs have contributed to creating. In contrast, pre-2022 printed books serve as fixed, human-created records that cannot be altered silently.
Another reason AI labs seek untouched paper is the ongoing struggle against data poisoning. Authors are beginning to push back by embedding elements in their texts that can be read by humans but not recognized by models, using techniques from tools like Nightshade. Research from Anthropic mentioned by ISBNdb highlights that as few as 250 to 500 meticulously crafted documents can introduce a backdoor into a dataset containing trillions of tokens.
Books published before 2022 evade this issue since they were authored prior to the introduction of such tools. ISBNdb markets this as provenance: by purchasing the physical book and keeping the paperwork, clients can ensure a clear legal chain of custody.
However, there is a challenge that the company acknowledges directly. Scanning books en masse typically results in their destruction. Workers cut off the spine so that loose pages can be fed into machines more quickly and cheaply than using careful methods. Thus, ISBNdb promises its clients confidentiality: “Strict NDA on every engagement,” states their website, making it clear that buyers’ identities are “never disclosed.”
The reasoning is also reflected in their marketing language. “The optics problem is real,” their site acknowledges. The headline “AI company destroys two million books” is unlikely to garner public sympathy, so one suggested alternative is to frame the action as a means of digitally preserving the works.
All of this intersects with an ongoing copyright dispute. A US judge recently approved Anthropic’s $1.5 billion settlement regarding pirated books and ruled that training on purchased, scanned books qualifies as fair use. The judge reasoned that the copying destroyed the original print copies, suggesting that one legal copy merely replaced another. Essentially, purchasing and shredding books is the legally endorsed method.
Consequently, the industry that vowed to digitize human knowledge is now buying and pulping it, load by load. Ingram, the largest book distributor in the US, has already alerted publishers and provided them with an option to opt out. It appears that the rarest commodity in AI is a sentence untouched by machines.
Other articles
AI companies are acquiring vintage books to ensure clean training data.
The broker ISBNdb is offering pre-2022 print books for AI labs as the final clean training data. Scanning these books involves discretely destroying them.
