Anna's Archive alleges AI firms scan and destroy books, calls for public preservation
- A guest post on Anna's Archive alleges that AI companies buy secondhand books through intermediaries, scan them for pre-2022 training data, and destroy the physical copies afterward.
- The post says Anthropic's confidential Project Panama, launched in early 2024 and described in a $1.5 billion copyright settlement, spent tens of millions of dollars on millions of paper books for Claude training.
- Anna's Archive argues that destructive scanning can leave AI companies as the only holders of digital copies, with the scans kept on private servers rather than available as readable source texts.
- The post asks volunteers to scan and upload books, journals, newspapers, magazines, ancient works, and rare materials, and says it can cover scanning costs and provide rewards for large uploads.
- It frames the effort as a race to digitize publications before publishers restrict access and before AI companies acquire and destroy more physical copies.
Hacker News opinions
If Anna's Archive scans everything, aren't AI companies just going to download it anyway? That seems like doing their work for them.
I would still rather have a public copy. The companies are not going to share the original scans.
I want the actual books to remain accessible, rather than their contents being mashed into a proprietary LLM and poorly regurgitated.
I see a real difference here. If a company destroys the only book after training, we lose direct access, provenance, citation, and the ability to pay for a specific text instead of metered inference.
I don't see the outrage over companies buying books and destructively scanning them. They own the copies, and most modern books exist in thousands or millions of identical copies.
We do not know what is being lost because the process is not transparent.
My concern is the smaller set of books with few surviving copies. Destroying one of those can eventually remove direct access altogether.
I think this view ignores how research works with old books. Some works are known only because other books cited them, and a rare surviving copy can preserve evidence about the past.
The concern is rare and out-of-print books, not mass-market copies. A recent academic book printed in 100 copies may be replaceable, but an 18th-century edition with one known surviving copy is a different case.
Buying a few potatoes and burning them is legal too, but buying enough food to feed a country and destroying it would still be wrong. Scale changes the ethical question.
The destruction is also powerful symbolism. It recalls Apple's 2024 iPad ad that crushed cultural objects into a glass slab, except this concerns actual books.
I agree it may be legal, but I find it disgusting. This is about what companies should do, not whether they are allowed to do it.
My main question is why the AI companies do not leak their scans to Anna's Archive. Are they really protecting a statistically meaningful training-data advantage?
I do not see why a profit-seeking company would give away an advantage to a free archive. That would hurt its own business.
Sharing the scans would also be illegal.
I was surprised and disappointed to hear Anthropic and other model companies may be doing this. They should preserve publicly accessible archives, even if many books have little training value, because pre-machine-generated texts have archival value.
They probably have high-quality scans, but the same copyright restrictions likely prevent them from publishing the archive.