AI companies are purchasing vintage books to acquire clean training data.
The message from the data broker ISBNdb is straightforward: the best AI training data in the world is available on bookshelves. The challenge lies in the fact that accessing this data requires removing the spines from millions of books and never revealing which laboratory financed the process.
AI faces a self-created pollution issue. An increasing portion of online content is now generated by machines, leading to a scenario where models begin to consume their own byproducts. One company believes that the solution is to utilize old, printed books released prior to the advent of chatbots.
According to a report by 404 Media, ISBNdb is now acquiring physical books in large quantities for AI laboratories to digitize for training data. The firm claims to be the largest book database globally. Its website states, “The world’s best AI training data is sitting on a shelf,” emphasizing that books are “dense, edited, authoritative."
The importance lies in the publication date. ISBNdb contends that books printed before 2022 are from the period before large language models emerged, so they wouldn't contain AI-generated content. This is significant due to the phenomenon of model collapse, which describes the documented decline in quality that occurs when models are trained on the artificial outputs of earlier models. Each iteration tends to be slightly inferior to the previous one.
The open web does not provide such certainty. A rapidly increasing portion of online text is now machine-generated, contributing to the same chaotic content cycle that the labs helped to create. In contrast, the fixed records of pre-2022 printed books are human-authored and cannot be quietly altered.
There is another motive for labs to seek unblemished printed material. Authors have begun to retaliate through data poisoning techniques, employing methods from tools like Nightshade to manipulate text in ways that are readable to humans but indecipherable to models. ISBNdb’s blog references research from Anthropic which indicates that as few as 250 to 500 specially crafted documents can insert a backdoor into a corpus containing trillions of tokens.
Books published before 2022, created prior to the existence of these tools, avoid this issue. ISBNdb promotes this as proof of authenticity: by purchasing the physical books and maintaining records, your legal team can ensure a clear chain of custody.
Nonetheless, there is a clear caveat that the company openly states. Scanning books at scale typically results in their destruction. Workers cut off the spines so loose pages can move through a machine, which is more efficient and cost-effective than the careful alternative. Thus, ISBNdb assures its clients of confidentiality. “Strict NDA on every engagement,” its website claims. The identities of buyers are “never disclosed.”
The reasoning behind this is evident in their marketing narrative. "The optics problem is real," reads the site. “'AI company destroys two million books' is not a headline that generates sympathy.” One proposed solution is to spin the destruction as a form of digital preservation.
All of this intersects with a copyright struggle that is already in progress. A US judge has approved Anthropic’s $1.5 billion settlement related to pirated books and determined that training on purchased, scanned texts qualifies as fair use. His rationale included the notion that the copying eliminated each original print, suggesting that one legal copy effectively replaced another. In essence, the process of buying and shredding books is the approach that courts will endorse.
Thus, the industry that once promised to digitize human knowledge now resorts to acquiring it and pulping it, one truckload at a time. Ingram, the largest book distributor in the US, has already alerted publishers and provided them with an option to opt out. It turns out the rarest commodity in AI is a sentence that has never been touched by a machine.
Other articles
AI companies are purchasing vintage books to acquire clean training data.
Broker ISBNdb is offering pre-2022 print books from AI labs as the final source of slop-free training data. The process of scanning these books involves quietly destroying them.
