AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

speckx1 pts0 comments

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

Account

Log in

Subscribe

Navigation

Home

About

RSS

Support/FAQ

Podcast

FOIA Forum Archive

Merch

Advertise

Privacy

Contact Us/Tips

Follow us

Twitter<br>Bluesky<br>Mastodon<br>Instagram<br>TikTok<br>Facebook<br>RSS

Advertisement

&bull;

Go ad free

News<br>AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop

Emanuel Maiberg

Jul 21, 2026<br>at 9:51 AM

ISBNdb, a company that sources printed books for AI companies to turn into training data, tells clients “the optics problem is real.”

Photo by Patrick Tomasso / Unsplash

As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.<br>“The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”<br>In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.<br>“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”<br>ISBN stands for International Standard Book Number, the numerical commercial book identifier and barcode on the back of most books. For years, ISBNdb helped book sellers, libraries, and distributors manage their inventory and find and sell books, but the generative AI boom has made it valuable to AI companies. In addition to selling access to book metadata, ISBNdb now helps AI labs source bulk printed book purchases of between 1,000 to 1 million books per order. ISBNdb’s data makes it easier for AI companies to methodically acquire, scan, and turn printed books into training data while avoiding duplication.<br>AI companies’ attempts to hoover up printed books for training data got wide attention in January after a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process. The Washington Post article found that Anthropic was buying books from one company called Better World Books, one of several marketplaces where libraries, retailers, and individuals can sell their books. Google was recently sued by book publishers for similarly training Google Gemini on copyrighted books.<br>ISBNdb advertises that it can keep the identity of AI companies secret.<br>“Strict NDA [non-disclosure agreement] on every engagement,” ISBNdb’s site says. “Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed.”<br>ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process.<br>“The optics problem is real,” ISBNdb’s site says. “‘AI company destroys two million books’ is not a headline that generates sympathy.”

One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. This bookseller asked to remain anonymous so he can continue to do business on these platforms.<br>“I personally have mixed feelings about all of this,” the bookseller, who suspects he’s sold hundreds of books to AI companies for training data, told me. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”<br>This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.<br>The seller told me that, normally, on a good week, he’d sell about 20 books. Since April, he has regularly sold hundreds of books a week....

books companies data isbndb training book

Related Articles