AI companies are purchasing vintage books to acquire clean training data.

AI companies are purchasing vintage books to acquire clean training data.

      The message from the data broker ISBNdb is straightforward: the best AI training data in the world is available on bookshelves. The challenge lies in the fact that accessing this data requires removing the spines from millions of books and never revealing which laboratory financed the process.

      AI faces a self-created pollution issue. An increasing portion of online content is now generated by machines, leading to a scenario where models begin to consume their own byproducts. One company believes that the solution is to utilize old, printed books released prior to the advent of chatbots.

      According to a report by 404 Media, ISBNdb is now acquiring physical books in large quantities for AI laboratories to digitize for training data. The firm claims to be the largest book database globally. Its website states, “The world’s best AI training data is sitting on a shelf,” emphasizing that books are “dense, edited, authoritative."

      The importance lies in the publication date. ISBNdb contends that books printed before 2022 are from the period before large language models emerged, so they wouldn't contain AI-generated content. This is significant due to the phenomenon of model collapse, which describes the documented decline in quality that occurs when models are trained on the artificial outputs of earlier models. Each iteration tends to be slightly inferior to the previous one.

      The open web does not provide such certainty. A rapidly increasing portion of online text is now machine-generated, contributing to the same chaotic content cycle that the labs helped to create. In contrast, the fixed records of pre-2022 printed books are human-authored and cannot be quietly altered.

      There is another motive for labs to seek unblemished printed material. Authors have begun to retaliate through data poisoning techniques, employing methods from tools like Nightshade to manipulate text in ways that are readable to humans but indecipherable to models. ISBNdb’s blog references research from Anthropic which indicates that as few as 250 to 500 specially crafted documents can insert a backdoor into a corpus containing trillions of tokens.

      Books published before 2022, created prior to the existence of these tools, avoid this issue. ISBNdb promotes this as proof of authenticity: by purchasing the physical books and maintaining records, your legal team can ensure a clear chain of custody.

      Nonetheless, there is a clear caveat that the company openly states. Scanning books at scale typically results in their destruction. Workers cut off the spines so loose pages can move through a machine, which is more efficient and cost-effective than the careful alternative. Thus, ISBNdb assures its clients of confidentiality. “Strict NDA on every engagement,” its website claims. The identities of buyers are “never disclosed.”

      The reasoning behind this is evident in their marketing narrative. "The optics problem is real," reads the site. “'AI company destroys two million books' is not a headline that generates sympathy.” One proposed solution is to spin the destruction as a form of digital preservation.

      All of this intersects with a copyright struggle that is already in progress. A US judge has approved Anthropic’s $1.5 billion settlement related to pirated books and determined that training on purchased, scanned texts qualifies as fair use. His rationale included the notion that the copying eliminated each original print, suggesting that one legal copy effectively replaced another. In essence, the process of buying and shredding books is the approach that courts will endorse.

      Thus, the industry that once promised to digitize human knowledge now resorts to acquiring it and pulping it, one truckload at a time. Ingram, the largest book distributor in the US, has already alerted publishers and provided them with an option to opt out. It turns out the rarest commodity in AI is a sentence that has never been touched by a machine.

Other articles

The Galaxy Z Fold 8 Ultra arrives featuring a more enhanced display, a larger battery, and a $100 increase in price. The Galaxy Z Fold 8 Ultra arrives featuring a more enhanced display, a larger battery, and a $100 increase in price. The new Galaxy Z Fold 8 Ultra from Samsung begins at a price of $2,099.99 and features a clearer display, larger battery, and enhanced ultrawide camera, all within a slightly thinner design. NEURA is constructing gyms to prepare robots for real-world applications. NEURA is constructing gyms to prepare robots for real-world applications. NEURA Robotics is launching a network of 'gyms' where robots can train in real-world conditions. Their belief is that the main limitation of Physical AI lies in experience rather than intelligence. Substack incorporates AI detection to combat ‘Claudefishing’. Substack has teamed up with the detector Pangram, allowing readers to check posts for the proportion of content created by humans compared to AI. This practice is referred to as Claudefishing. Revolut reaches a valuation of $115 billion and receives a banking license in Australia. Revolut reaches a valuation of $115 billion and receives a banking license in Australia. A secondary share sale has valued Revolut at $115 billion, which is a 53% increase since November, as it secures the first Australian banking license for a global fintech. Samsung announces the Z Fold8 Ultra, Z Flip8, Watch Ultra2, and AI eyewear during Galaxy Unpacked. Samsung announces the Z Fold8 Ultra, Z Flip8, Watch Ultra2, and AI eyewear during Galaxy Unpacked. Samsung unveils six products at the Galaxy Unpacked event in London, featuring its inaugural Ultra foldable, an upgraded Fold8, and AI eyewear in collaboration with Gentle Monster. AppLovin Supports Velocity As The AI Sector Searches For Growth Beyond Subscriptions AppLovin Supports Velocity As The AI Sector Searches For Growth Beyond Subscriptions Velocity, supported by AppLovin and a $27 million seed funding round, is developing ad-monetization infrastructure for AI applications, claiming that intent-based advertising can provide funding for free access when subscriptions are inadequate.

AI companies are purchasing vintage books to acquire clean training data.

Broker ISBNdb is offering pre-2022 print books from AI labs as the final source of slop-free training data. The process of scanning these books involves quietly destroying them.