Bulk secondhand book orders raise questions about AI firms' destructive scanning
- A 2025 US ruling in authors' case against Anthropic found that training AI on books it had purchased was not copyright infringement, calling the use "exceedingly transformative."
- Unsealed court documents described Anthropic's Project Panama goal as "destructively scan all the books in the world," a process that removes spines for industrial scanning and recycles the remains.
- Independent sellers report unexplained bulk orders: Barter Books normally sells 2,000 to 3,000 books per week, but a single Canadian-company order recently matched a week's sales.
- Anthropic says Claude uses public web data, commercially acquired datasets and internally generated data, and says none of its acquisition programs buy and destroy rare or antiquarian books.
- Oxford intellectual-property professor Emily Hudson says UK law differs from the US: copying to build a training library and to train a model generally requires the copyright owner's permission.
Hacker News opinions
New books seem more expensive while editing and print quality decline, so I buy used books because it makes financial sense. I am not convinced AI explains the wider used-book market.
I buy almost entirely used now. My local independent shop joined a used-book platform: I tell them what I want to sell, bring it in when they find a buyer, and they take a handling fee. Used copies are often half the price of new.
In the US, anything without assured mass-market demand is likely print-on-demand. Some print-on-demand books are fine, but conventional printing is often better quality.
I prefer mass-market paperbacks, but publishers seem to have dropped them because they cannot charge more than $20. I can buy decades of good books for under $1, and physical books do not track me or show me ads.
The bulk orders make AI training the likely cause. Buying the books pays the party that owns the copy, and training seems at least as transformative as putting text in a web-search index.
My objection is not copyright. Mass-buying books and destroying them by the millions is offensive, especially when rare copies may disappear. Paying for them does not settle that.
US law creates a wasteful result: if only 25 copies exist, 25 AI companies may each buy and pulp one. Scan the books, preserve the files, compress and back them up, then make the archive widely available.
Copyright is broken for out-of-print books. If a publisher will not sell a new copy, I do not see who loses money when someone copies it.
The judge did not specifically require destruction. They likely cut off the spines because destructive scanning is faster.
I do not think destroying the physical copy is automatically wrong. An artist's value depends partly on copies in circulation, but copyright should expire far sooner, perhaps after 13 years.
I donate to Anna's Archive because preserving and distributing books matters. It supports more than 20 payment methods.
AI companies can afford bulk book purchases and data centres with investment money. Anthropic reportedly projected at least $10.9 billion in Q2 2026 revenue and a first quarterly operating profit of $559 million.