AI companies destroy physical books – let's scan rare books before it's too late
AI companies are bulk-buying secondhand books, scanning them for training data, and then destroying the physical copies, citing U.S. copyright rules that favor “format shifting” when only one copy exists. Critics see this as a privatization and potential loss of cultural heritage—especially for out‑of‑print or poorly archived works—arguing that destruction should at least be paired with public access to the scans. Others counter that most of these titles are low‑demand manuals and vanity press books that would be pulped by libraries or resellers anyway, and that the real problem lies in copyright law and policy rather than the scanning itself.
What AI companies are actually doing
- Several AI labs are buying large quantities of second‑hand books (often via intermediaries), cutting off the spines, scanning pages, then discarding or pulping the paper.
- Motives discussed:
- Mainstream claim: to obtain high‑quality, pre‑AI (“untainted”) text and to comply with recent US copyright rulings allowing “format‑shifting” if the original is destroyed.
- Skeptical view: destruction is mostly about cost, logistics, and maintaining a proprietary data moat, not legal necessity.
Copyright and legal context
- Recent US cases (e.g., Project Panama / Bartz v. Anthropic as paraphrased in the thread):
- Scanning purchased physical books and destroying them can be treated as a single transformed copy, favoring a fair‑use argument.
- Keeping both physical and digital copies, or using pirated ebooks, runs higher infringement risk.
- Many commenters stress: the law doesn’t literally require shredding; it just makes destructive scanning the safest option.
- Strong criticism of current copyright: long terms, orphan works, out‑of‑print but still locked content; blame placed on both legislators and publishers.
Are “rare books” really being lost?
- One camp: most scanned books are low‑demand items—old manuals, obscure conference proceedings, vanity press, TV guides—likely headed for dumpsters or pulpers anyway.
- Other camp: at this scale, some truly rare or unique works (small academic runs, local histories, niche cultural artifacts) will inevitably be destroyed.
- Evidence is thin: discussion repeatedly notes a lack of concrete examples of specific unique titles lost; many claims are hypothetical or anecdotal.
Libraries, booksellers, and scale
- Multiple commenters note that:
- Public libraries, university libraries, thrift stores, and publishers routinely discard or pulp huge numbers of books.
- Donation books often go straight to recycling; “last copy” policies and professional archivists are the exception, not the norm.
- Others push back:
- Libraries at least have some archival ethos and trained staff; AI companies have profit motives, minimal curation, and much larger, targeted throughput.
- Intention and scale (millions of deliberately purchased books) are seen as morally relevant differences.
Digitization vs. public access
- Key tension: AI labs keep high‑quality scans private for legal and competitive reasons; the public only sees lossy, paraphrased traces via models.
- Many argue that:
- Training on a book is not equivalent to preserving it; LLMs can’t reliably reproduce exact text, citations, or context.
- Scans locked on corporate servers are “barely better than landfill” unless laws change to allow eventual public release.
- Proposed fixes: mandatory deposit of digital copies with national libraries, automatic public release after copyright expiry or for out‑of‑print works, or explicit legal carve‑outs for public archives.
Values, rhetoric, and proposed responses
- Strong emotional reactions:
- Some liken destructive scanning to book burning or a modern Library of Alexandria moment; others call this hyperbolic “moral panic.”
- Counter‑view: this is just another stage in the normal life cycle of surplus books, with the side benefit of digital preservation.
- Suggested actions:
- Support and volunteer scanning for shadow libraries (Anna’s Archive, LibGen) and the Internet Archive.
- Push for copyright reform and library‑style exceptions; some call for boycotts of large AI companies.