AI companies are acquiring vintage books to ensure clean training data.

AI companies are acquiring vintage books to ensure clean training data.

      The message from data broker ISBNdb is straightforward: the best AI training data available is currently sitting idle on shelves. The downside is that accessing this data requires cutting the spines off millions of books, and the identity of the paying labs must remain undisclosed.

      AI is facing a self-created pollution issue. Much of the web now consists of machine-generated content, leading to models potentially consuming their own byproducts. One company believes the solution lies in old printed books released before the chatbot era began.

      404 Media reported that ISBNdb is now acquiring physical books in bulk for AI labs to utilize as training data. The company claims to be the largest book database globally. Their website asserts that “the world’s best AI training data is sitting on a shelf,” emphasizing that books are “dense, edited, and authoritative.”

      The significance is in the publication date. ISBNdb asserts that books printed before 2022 were produced before the explosion of large language models, indicating they lack AI-generated text. This distinction is crucial due to model collapse, which refers to the decline seen when models are trained on the outputs of previous models, resulting in each generation being slightly inferior.

      The open web cannot provide such assurance. A rapidly increasing portion of online content is machine-generated, part of the same mess that AI labs have contributed to creating. In contrast, pre-2022 printed books serve as fixed, human-created records that cannot be altered silently.

      Another reason AI labs seek untouched paper is the ongoing struggle against data poisoning. Authors are beginning to push back by embedding elements in their texts that can be read by humans but not recognized by models, using techniques from tools like Nightshade. Research from Anthropic mentioned by ISBNdb highlights that as few as 250 to 500 meticulously crafted documents can introduce a backdoor into a dataset containing trillions of tokens.

      Books published before 2022 evade this issue since they were authored prior to the introduction of such tools. ISBNdb markets this as provenance: by purchasing the physical book and keeping the paperwork, clients can ensure a clear legal chain of custody.

      However, there is a challenge that the company acknowledges directly. Scanning books en masse typically results in their destruction. Workers cut off the spine so that loose pages can be fed into machines more quickly and cheaply than using careful methods. Thus, ISBNdb promises its clients confidentiality: “Strict NDA on every engagement,” states their website, making it clear that buyers’ identities are “never disclosed.”

      The reasoning is also reflected in their marketing language. “The optics problem is real,” their site acknowledges. The headline “AI company destroys two million books” is unlikely to garner public sympathy, so one suggested alternative is to frame the action as a means of digitally preserving the works.

      All of this intersects with an ongoing copyright dispute. A US judge recently approved Anthropic’s $1.5 billion settlement regarding pirated books and ruled that training on purchased, scanned books qualifies as fair use. The judge reasoned that the copying destroyed the original print copies, suggesting that one legal copy merely replaced another. Essentially, purchasing and shredding books is the legally endorsed method.

      Consequently, the industry that vowed to digitize human knowledge is now buying and pulping it, load by load. Ingram, the largest book distributor in the US, has already alerted publishers and provided them with an option to opt out. It appears that the rarest commodity in AI is a sentence untouched by machines.

Other articles

Cisco's small open-weight AI detects bugs, outperforming Gemini. Cisco’s open-weight Antares models identify code bugs on your machines. The company asserts that they outperform Gemini and GPT-class systems, being faster and significantly more cost-effective. Galaxy Z Fold 8 vs. Galaxy Z Fold 8 Ultra: Is it worth spending an additional $200 for the Ultra model? Galaxy Z Fold 8 vs. Galaxy Z Fold 8 Ultra: Is it worth spending an additional $200 for the Ultra model? The Samsung Fold 8 and Fold 8 Ultra utilize the same chip and storage options, but they differ significantly in terms of camera hardware, screen size, and battery. 9 top requirements management tools for the development of medical devices (2026) 9 top requirements management tools for the development of medical devices (2026) Nine tools for managing requirements in medical device development, evaluated based on traceability, support for compliance, and readiness for FDA and EU MDR regulations. The Galaxy Watch 9 features a new processor, larger batteries, One UI 9 Watch, and a $30 increase in price. The Galaxy Watch 9 features a new processor, larger batteries, One UI 9 Watch, and a $30 increase in price. The Galaxy Watch 9 replaces Samsung's Exynos chip with the Snapdragon Wear Elite, introduces FDA-approved health features, and has a starting price of $379.99, which is an increase of $30 compared to the Watch 8. In June, electric vehicle sales in Europe skyrocketed by almost 40 percent, with market share exceeding 25 percent. In June, electric vehicle sales in Europe skyrocketed by almost 40 percent, with market share exceeding 25 percent. In June, battery-electric vehicle registrations reached 275,060 in 17 European countries, leading to first-half sales surpassing 1,240,000 and a market share exceeding 25 percent. Instagram is finally providing a soundtrack revamp for your old posts. Instagram is finally providing a soundtrack revamp for your old posts. Instagram's new Replace Audio feature lets users change the music on already posted feed posts and carousels while maintaining their engagement.

AI companies are acquiring vintage books to ensure clean training data.

The broker ISBNdb is offering pre-2022 print books for AI labs as the final clean training data. Scanning these books involves discretely destroying them.