AI companies are buying and then destroying millions of rare books to prevent AI slop content that consumers are vocally against, 404 Media reported last week.
Silicon Valley tech giants are paying companies and contractors to buy up rare books, which are then scanned in a high-speed machine that cuts their spines out and then shreds the originals.
The tech giants are reportedly buying up the rare books to train new AI models and prevent AI “slop.” In one article on its site, ISBNdb, a company that claims to have the “the world’s largest book database,” argued that books published before 2022 were best for AI training data because they would not have any AI-generated text.
In a report uncovered by the Washington Post in January, one Anthropic co-founder suggested that feeding AI models books could teach them “how to write well” instead of producing “low-quality internet speak.”
“The world's best AI training data is sitting on a shelf,” ISBNdb wrote in a since-deleted blog post. “Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”
404 Media reported that the AI tech companies were interested in buying up books to prevent “model collapse,” where AI models trained on lower-quality AI-generated results progressively lose quality. Executives believed that vast troves of books were essential to give the models new data and prevent their model collapse.
Why are AI companies destroying old books?
Notably, special edition book sellers say that they usually sell one or two books to a single customer.
“In the rare book trade, it’s very seldom that people want to buy more than one book,” antique book seller Pieter de Vries told The Telegraph in an interview published last week. “So if somebody comes and says, ‘I want a couple of hundred of your books,’ it’s very strange.”
But ISBNdb and companies like it are now helping AI tech giants purchase orders ranging from 1,000 to one million books.
The AI companies’ attempt to hoover up books, art, news articles, and other forms of media has gotten attention in several instances since 2024.
Most notably, the Washington Post reported in January on Anthropic’s attempts to buy millions of books, slice their spines, and scan their pages to feed more data into the company’s chatbot, Claude.
Anthropic's legal battle over AI and book copyrights
That instance later led to a multi-million dollar class action lawsuit in which several authors sued Anthropic, arguing that the company, which is backed by Amazon and Alphabet (Google's parent company), used pirated versions of their books without permission to teach Claude to respond to human prompts.
Anthropic settled with the authors at $1.5 billion.
Notably, though, the fact that Anthropic destroyed the books made the company’s case stronger.
Judge William Alsup ruled last June that Anthropic made fair use of the authors' work to train Claude, but found that the company violated their rights by saving more than seven million pirated books to a "central library" that would not necessarily be used for AI training.
“Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,” Alsup wrote.
“The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company,” he said.
In short, because the AI company pulped the books, the digital version that existed afterward “replaced” the physical one. The judge ruled that the companies merely transformed the books, which meant that it was not a violation of US copyright law.
The tech giants clearly wanted this secret since the beginning, because the optics of destroying books are unpalatable for many. Documents uncovered by the Washington Post found that even internally, companies wanted to distance themselves from the effort.
“Project Panama is our effort to destructively scan all the books in the world,” one internal document unsealed in legal filings last said, as reported by the Washington Post. “We don’t want it to be known that we are working on this.”
As ISBNdb wrote in a since-deleted blog post on its website: “The optics problem is real. ‘AI company destroys two million books’ is not a headline that generates sympathy.”
Tech giants, former officials decry AI companies for destroying books
But the headlines and lawsuits became public, and the public is rather unsympathetic.
“So AI labs are buying old books by the pallet, slicing them apart, scanning the pages, and pulping what's left,” former US House rep. Brad Carson wrote in a post on X/Twitter.
"Here's the perverse part. A federal court blessed this precisely because the original is destroyed. One legal copy replaces another, so it's fair use. Whatever you think of that ruling or fair use, notice what it does. The law now rewards destruction and penalizes preservation.
“A lab that wants to scan a book and keep it, or donate it, or deposit the scan in a public archive, has weaker legal footing than a lab that shreds everything. We have built a legal machine that pays people to pulp books and punishes them for saving them.”
Some tech giants have voiced their distaste for the destruction of the books.
“I’ve asked the SpaceX AI team to preserve any rare books in a library and scan them the hard way rather than just cutting off the spine and scanning,” Elon Musk said in an X/Twitter post.
"There's something particularly misanthropic about the mechanized destruction of such intimate human objects," said SEO of Factory AI Matan Grinberg. "History seldom looks kindly on those who destroy books, whatever the reasons."
One critic of the AI industry’s approach to copyrighted work and the founder of Fairly Trained, a creator-rights group, Ed Newton-Rex, told the Telegraph that the secrecy with which companies like ISBNdb and Anthropic acquire books is damning.
“Clearly both the provider of these books and the AI companies know that this is a terrible look and they don’t want the specifics to get out,” he told the Telegraph.
“If you are just going and spending $1 on a used book, with all of the money going to a book wholesaler, should that give you the right to train a commercial generative AI model on that book, which will then be able to compete with the author who wrote it?” Newton-Rex said. “A lot of people, myself included, think it shouldn’t.”
“There is surely no more fitting image in the generative AI age for the exploitation that underlies this technology than almost trillion-dollar companies buying books for a few cents or a dollar each, scanning them, training on them, then destroying them, essentially subsuming culture,” he added.
In response to the flurry of reporting around the destroyed books, ISBNdb disputed the reports that it helped purchase large volumes of books for AI training.
“We've seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised,” the company wrote in a statement.
“The facts: ISBNdb has never purchased, scanned, or sold a book - for AI training or anything else. We don't train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We've taken the page down.
“Our job is helping people find books. For more than two decades, ISBNdb has been the card catalog of the book world - the data behind how bookstores, libraries, and reading apps connect readers with titles. Data about books, not the books themselves. That hasn't changed.”