Major AI companies are allegedly purchasing and destroying millions of physical books through third-party intermediaries to avoid public backlash regarding their data harvesting practices, according to a report from 404 Media. This practice comes as the demand for unique, human-authored text for training large language models has increased, with physical books predating AI-generated content being particularly valuable.
Internal documents, unsealed during a lawsuit against AI firm Anthropic, revealed the company's awareness that destroying millions of books for training its Claude models, even if deemed transformative under fair use, could generate negative publicity. "We don’t want it to be known that we are working on this," Anthropic stated in planning documents, seeking to prevent its chatbot from producing output resembling "low quality internet speak."
To circumvent potential PR crises, AI providers are reportedly engaging services like ISBNdb, which until recently offered book sourcing specifically for "LLM training needs." An archived version of ISBNdb's site boasted the ability to procure up to one million books per order from various sources, including used bookstores and out-of-print catalogs. The service also promised strict non-disclosure agreements to protect client identities and acquisition targets, noting that "'AI company destroys two million books' is not a headline that generates sympathy."
ISBNdb has since removed the page detailing its AI book-sourcing service, stating it was part of exploring demand and that they have chosen to pivot away from that direction. However, booksellers specializing in rare and low-circulation titles report an unprecedented surge in sales since April, with some fulfilling hundreds of book orders weekly. These bulk orders often consist of seemingly random books, all possessing ISBNs, suggesting they are identified through ISBN-based databases.
Booksellers have expressed mixed feelings about the trend. While the sales are financially beneficial and help clear out old inventory, some harbor reservations about the end-use of the books. The practice involves scanning the books, de-spining them, and ultimately discarding them, a process that, in Anthropic's case, was ruled by a judge to constitute fair use under copyright law due to its transformative nature. Despite this legal precedent, the desire to operate anonymously highlights the AI industry's sensitivity to public perception regarding its data acquisition methods.