mrkeyoor.com_
Sun 16 Aug 07:28 UTC
AI16 Aug 2026 05:37 UTC6 min read

AI’s Search for Data Leads to Secondhand Bookstores

Mysterious bulk purchases are clearing shelves at used bookstores. The buyers are suspected to be AI firms, turning physical books into training data and pulping the remains.

Secondhand booksellers are reporting a strange new type of customer: buyers who want everything, pay instantly, and don’t care about condition, genre, or title. These mysterious bulk orders are clearing out stock from charity shops and independent sellers, and the primary suspect is the artificial intelligence industry. According to reports, AI companies, hungry for the vast amounts of text needed to train Large Language Models (LLMs), may be turning to the analog world of physical books as a new source of data. The practice highlights a critical bottleneck in AI development—the finite supply of high-quality digital text—and raises new questions about copyright, market disruption, and the ultimate fate of the printed word.

Reports from booksellers describe a consistent pattern. Buyers, often operating through intermediaries, place orders for thousands of books at a time. They are largely indifferent to the content, requesting only that the books be in English and readable. This has led to speculation, now widely discussed among booksellers, that the books are not for reading but for processing. The BBC reports that booksellers believe the books are being scanned to digitize their text, and then pulped. The book is not the product; the data inside it is.

This trend appears to be widespread. In the UK and Ireland, booksellers have voiced similar suspicions about what they call ‘strange’ bulk orders, a story that gained significant traction on social media platforms like Mastodon. The phenomenon points to a brute-force solution to a digital problem. While the internet provided the first massive trove of data for training AI, that resource is being exhausted. AI firms have already scraped vast portions of the public web, and now face diminishing returns and a growing number of lawsuits from publishers and creators over copyright infringement.

The Insatiable Appetite for Data

To understand why an AI firm would bother with the logistics of buying and scanning millions of physical books, one must understand the mechanics of LLMs. These models learn language, context, and reasoning by analyzing statistical patterns in enormous text datasets. The quality and diversity of this training data directly determine the model's capability. Initially, developers used freely available text from sources like Wikipedia, Common Crawl (an archive of the web), and digitized public domain books.

As models grew larger and more sophisticated, their data requirements exploded. Researchers now speak of a “data wall,” a point at which the supply of high-quality, easily accessible digital text will be exhausted. To continue improving their models, companies need new, untapped sources. Books represent an ideal, if challenging, reservoir. They contain high-quality, well-edited, long-form text spanning every conceivable topic. Much of this material, especially from older and out-of-print books, has never been properly digitized or is locked behind publisher paywalls.

Acquiring physical books offers a few potential advantages for an AI company. First, it sidesteps the licensing complexities and costs of dealing directly with publishers for their digital catalogs. Second, it provides access to a unique dataset that competitors who rely solely on web scraping might not have. Third, it operates in a legal gray area. While the text in a book is copyrighted, the physical object itself is subject to the first-sale doctrine, which allows the owner of a copy to sell or dispose of it freely. An AI firm could argue that scanning a book it physically owns for internal data processing constitutes fair use—a claim that is legally untested and highly contentious, but one that avoids the clear copyright violations alleged in web-scraping lawsuits.

A New Front in the Copyright Wars

The move from scraping digital content to acquiring physical media marks a new phase in the ongoing conflict between AI developers and copyright holders. Major AI labs, including OpenAI, Google, and Meta, are already facing numerous lawsuits from authors and publishers who allege their work was used without permission to train models like ChatGPT and Llama.

These lawsuits focus on the unauthorized copying of digital files. The act of buying a physical book, scanning it, and using the resulting text for training purposes introduces a new set of legal arguments. AI companies could frame the process as a form of “reverse engineering” a text they legally own to extract non-expressive data—the statistical patterns of language—rather than reproducing the creative work itself. Publishers and authors, in contrast, will likely argue that this is simply a more laborious method of making an unauthorized digital copy for commercial gain, which constitutes clear infringement.

The Guardian’s report notes that this development comes after it was revealed that the AI firm Anthropic had spent millions on books for “data acquisition.” While the specifics of that program are not public, it lends credibility to the theory that major, well-funded players are actively pursuing this strategy. It suggests a calculated decision to invest in the significant logistical overhead of physical book processing as a potential workaround to the legal and practical limits of digital data sourcing.

Market Disruption and Collateral Damage

While AI firms may see secondhand books as a raw commodity, their actions have tangible consequences for the ecosystem they are entering. For many small booksellers and charity shops, these bulk orders have created a short-term financial boom. It’s a tempting offer: a single buyer willing to clear out slow-moving inventory with no haggling.

However, this influx of industrial-scale purchasing distorts the market. It can drive up the price of bulk used books, making it harder for smaller sellers, collectors, and readers on a budget to find affordable stock. The very nature of the transaction—where books are treated as interchangeable units of text—removes them from circulation permanently. A book that is scanned and pulped is a book that can never be read by another person, studied by a researcher, or discovered by a collector.

This raises concerns about the preservation of cultural heritage. Among the indiscriminately purchased pallets of paperbacks and hardcovers may be rare, out-of-print, or first-edition copies that hold cultural value beyond their text. When the goal is simply to feed a machine, these distinctions are lost. The book as a physical artifact, a piece of history, is rendered irrelevant. The process is one of consumption, turning a durable cultural good into a disposable input for a computational process.

The Logistics of Analog-to-Digital Conversion

The scale of this operation should not be underestimated. It requires a significant industrial pipeline. After purchase, the books must be transported to a central facility. There, they would likely be processed by high-speed, destructive scanners that cut off the spines to quickly feed individual pages. Following the physical scan, Optical Character Recognition (OCR) software converts the page images into machine-readable text.

This OCR process is itself a challenge. Old books with unusual fonts, yellowed paper, or annotations present difficulties for even the best software. The resulting raw text would be “noisy” and require extensive cleaning and pre-processing before it could be used for training an LLM. This entire chain—from sourcing and logistics to scanning and data cleaning—requires a level of capital and technical expertise that points directly toward large, well-funded technology companies.

What to Watch Next

The suspected entry of AI firms into the secondhand book market is a developing story, but it signals a broader shift in the hunt for training data. As the low-hanging fruit of the public internet is plucked, the search for proprietary, high-quality data will intensify, pushing companies into more unconventional and legally ambiguous territory.

Moving forward, the key area to watch will be the legal response. It is almost certain that authors’ guilds and publishing houses will challenge this practice as a form of mass copyright infringement. The outcome of these challenges could set a crucial precedent for whether the contents of a physical object one owns can be freely converted into data for commercial AI training. Furthermore, the industry will face growing pressure for transparency. Demands for AI companies to disclose their training data sources will likely grow louder, as the impact of their acquisition methods becomes more visible in markets far removed from Silicon Valley.

Finally, the market itself will adapt. Some booksellers may choose to refuse anonymous bulk sales, while others may embrace the new revenue stream. This could lead to a bifurcation in the market, with one stream feeding industrial data pipelines and another serving traditional readers and collectors. What is clear is that the insatiable data appetite of AI is no longer a purely digital phenomenon; it has crossed into the physical world, and the shelves of your local bookstore are its new frontier.

Sources

  1. Secondhand book sales are booming. Is it because of AI?
  2. Secondhand booksellers in UK and Ireland suspect AI firms behind ‘strange’ bulk orders