BenchmarksNewsPC Components

AI trains on shredded books tech giants secretly buy them up

Featured image AI trains on shredded books tech giants secretly buy them up

The race for advanced artificial intelligence is hitting a new, unexpected obstacle: a desperate hunger for truly original, high-quality information. As AI models continue to grow exponentially, their reliance on massive datasets has brought the quality of that training material into sharp focus, pushing tech giants to look far beyond the internet for knowledge.

The problem is simple: if AI is trained on mediocre content—so-called “AI slop”—the resulting intelligence will be inherently flawed. To achieve true cognitive depth, leading AI developers are actively seeking human-authored sources that predate the current digital deluge, turning their attention toward humanity’s literary history.

This quest has led some of the most influential players in the field to explore a radical solution: physical books. It is not just an academic curiosity; it is becoming a strategic move for data sourcing.

For instance, AI entities have already experimented with this approach. One major player, Anthropic, reportedly invested millions into extracting knowledge from countless printed volumes to build its advanced Claude AI models. This effort sparked significant legal battles, highlighting the thorny intersection of copyright law and machine learning. While courts have generally affirmed that using books for training falls under fair use, the scale of these operations has resulted in staggering settlements, demonstrating how fiercely intellectual property is protected in the age of automation.

Meanwhile, secondary platforms designed to organize this massive influx of information are capitalizing on the trend. Online databases like ISBNdb, which catalog over 111 million books, have pivoted their focus. They are now offering specialized services to bulk-purchase books for AI companies, facilitating transactions ranging from a few thousand copies to as many as one million in a single deal.

This surge in demand has dramatically altered the market. Professional booksellers and distributors have reported unprecedented sales spikes, with weekly volumes increasing fivefold, signaling that physical texts are rapidly transitioning from mere consumption to a vital commodity for the AI economy.

The process of digitalization is complex. Some methods involve high-speed industrial scanners, while others utilize more hands-on techniques like scanning or even destructive methods to feed individual pages into machines, prioritizing efficiency and cost reduction over preservation.

However, this pursuit comes with profound ethical questions. While the potential for smarter AI is enormous, there remains a debate about the ethics of removing physical books from circulation. The challenge lies in ensuring that these digitalized treasures are not filtered out—especially rare or out-of-print titles—and that the resulting massive datasets do not diminish access to knowledge for future generations.

The ultimate prize is smarter AI, but the cost of this revolution may be measured not just in billions of dollars, but in the preservation and accessibility of human thought itself.