The Devouring of All Knowledge - Ryan FleckThe Devouring of All Knowledge
1,171 words<br>6 min. read<br>Pub.<br>Jul 2026<br>Upd.<br>Jul 2026<br>I was horrified but unsurprised to read reports1 this week of the<br>AI/LLM2 companies buying, butchering, scanning, and shredding rare<br>books at an unprecedented scale.<br>With all online forums, PDFs, articles, and books consumed, the next<br>best source of human writing is the abundance of physical books<br>published before the advent of LLMs. The year 20213 and the<br>emergence of LLM2 text generation is coarsely termed “the<br>ensh**tification of all knowledge”. Writing published after this<br>point is potentially LLM-written and therefore unsuitable to be used<br>as training data. Whatever we wrote before this historic point is<br>exclusively the work of human hands and minds.<br>The race is now on to consume and destroy (at least one copy of) every<br>physical book on the planet first - and I wouldn’t rule out a scenario<br>similar to the global memory shortage of 2025. At this time, purchases<br>became weapons used to strangle the supply chains of competitors and<br>gain an exclusive material advantage. Memory prices shot far out of<br>reach for normal people as AI/LLM companies rushed to secure the<br>remaining global supply of DRAM chips. Will we see the same thing<br>happen with rare, old, and out-of-print books?<br>X users commenting on the book-hoarding situation.<br>“Many companies, including Anthropic, have turned to ingesting<br>physical books instead, which they can buy countless used copies of<br>on the cheap. According to the settled lawsuit, Anthropic used a<br>hydraulic powered cutting machine to neatly remove the pages from<br>the books it procured from book resellers and then scanned them<br>using industrial-grade imaging equipment. In other words, it was<br>literally ripping off authors’ books to train its AI.”<br>– Frank L. Futurism1
The problem here is not just the destruction of rare and priceless<br>books which, though currently unused, are presently accessible human<br>knowledge - hoarding this rare knowledge behind walled gardens is<br>the real issue. Countless dystopian sci-fi visions of humanity’s<br>darkest paths describe the gathering and destruction of knowledge so<br>the record can be set straight by a central authority. Ray Bradbury’s<br>Fahrenheit 451 and George Orwell’s Nineteen Eighty-Four both come<br>to mind as keen examples of governments and their proxies working<br>tirelessly to destroy the past and rewrite history.<br>Ultimately, this leads to a nightmare situation where all knowledge is gathered<br>by the tech companies feeding these insatiable large language models.<br>Every copy of every book will be frantically secured to keep them out<br>of the hands of competitors. Knowledge will then become accessible<br>exclusively through the censored and monitored outputs from these<br>models. Information deemed ‘unsafe’ will be hidden forever underneath<br>the safety, compliance, and ethics layers of the largest models -<br>governing and correcting the output of every individual response.<br>Whatever is deemed harmful by regulators will be unavailable for the<br>greater good. History will be gatekept.<br>An argument could be made that these books are much better put to use<br>in this manner. By digitizing and using the books as training data,<br>the knowledge is made more accessible to the everyman who uses<br>ChatGPT as their primary knowledge source. I would be more excited<br>about this if the scans and content were not to be permanently<br>sequestered in each LLM company’s knowledge silo. No guarantee is<br>provided that any of this material will be presented in an impartial<br>manner, or ever released to the public - this would remove the<br>expensive knowledge advantage the particular LLM company gained<br>through the acquisition and scanning process.<br>“One small book seller said that in April, he suddenly went from<br>selling no more than 20 books a week to hundreds , and he’s almost<br>certain that the customers are AI labs, noting the random selection<br>of the books and how they all have ISBNs . He added that his<br>inventory is full with rare and out-of-print books, meaning that an<br>AI company could be destroying some of the few remaining copies that<br>can be found.”<br>– Frank L. - Futurism1
“Obviously, printed books represent a treasure trove of information<br>for AI. However, many debate the ethics of removing books from<br>circulation since it is uncertain whether AI companies filter the<br>rare or even out-of-print books from the common titles during<br>digitalization. The other major issue is that scanned books go<br>directly into a private database to train AI, which the general<br>public does not have access to. True, we will have smarter AI, but<br>at the cost of the information not being available to future<br>generations.”<br>– Zhiye L. - Tom’s Hardware4
What can you do about it? Learn to read, then buy and hold paper<br>books. Your children will thank you, and paper books are entirely<br>yours to mark up, borrow, trade, and love -...