AI companies are shredding rare books
AI companies are reportedly bulk-buying and destructively scanning print books to create training data for large language models, then shredding the physical copies to strengthen their legal claim that only “one copy” exists. Commenters argue over whether this amounts to modern book burning or pragmatic format-shifting, questioning how “rare” the affected titles really are, and pointing out that copyright law and DMCA anti-circumvention rules make non-destructive scanning and public release of scans legally risky. Many see the core problem not in digitization itself but in a regime that incentivizes destroying physical media, hoarding the resulting datasets, and sidelining public institutions like libraries and archives.
Scope and evidence of shredding
- Thread centers on reports (esp. a 404 Media piece) that AI-adjacent firms bulk-buy used books, scan them destructively, then pulp the remains.
- Concrete examples cited are mostly niche, recent technical or academic titles with ISBNs (post‑1960s), not ancient manuscripts.
- Claims about “last copies” of 18th‑century works being shredded are acknowledged by several commenters as hypothetical, not evidenced.
- Some participants doubt any truly rare/out‑of‑copyright books are being destroyed; others say that with foreign‑language and low‑circulation works, single‑digit known copies are plausible but unproven here.
Copyright, legality, and incentives
- Several note U.S. court decisions finding that training on lawfully acquired works can be fair use, especially when the physical copy is destroyed so only one copy exists.
- Destructive scanning is described as faster, cheaper, and legally safer than keeping both physical and digital copies under current copyright/DMCA rules.
- Some blame publishers and past lawsuits (e.g., against Internet Archive, shadow libraries) for pushing AI firms toward “analog hole” strategies instead of licensing.
- Others emphasize that if books are out of copyright, there is no legal need to shred them; if it happens, it’s for cost/speed, not law.
Preservation vs. destruction
- One camp argues this is effectively book burning: unique artifacts (bindings, marginalia, palimpsest layers) and physical heritage are lost, while private scans are locked away.
- Another camp counters that most of these books are already valueless surplus that libraries and publishers routinely weed and pulp; destructive scanning at least preserves content digitally.
- Some stress that LLM weights are not proper digitization; if scans are not publicly accessible, nothing is truly preserved for scholarship.
- Others think physical “fetishization” is overblown and that the primary value is the text, not the paper, though digital longevity and authenticity are contested.
Comparisons, ethics, and analogies
- Comparisons range from Nazi/Taliban destruction and Fahrenheit 451/“book burning as a service” to mundane library weeding and publisher pulping.
- Several note it is possible for everyone in this conflict—publishers, AI companies, lawmakers—to be “wrong” in different ways.
- Some suggest AI firms should be required (or socially pressured) to release scans of public‑domain books, or to contribute to open archives, as a condition of destructive scanning.
Policy and reform ideas
- Repeated calls for copyright reform: shorter terms (e.g., 20–50 years), stronger public‑domain and abandonment rules, and clearer rights for non‑destructive digitization.
- Proposals include public or nonprofit custodians for rare works, patrimonial protection for single‑digit‑copy books, and public‑owned training corpora instead of private hoards.