We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility

Amazon has been traced buying large batches of “rare” used books that end up at facilities where they are scanned for AI training and the physical copies are reportedly destroyed. Commenters debate whether this is cultural vandalism or overblown outrage given that many such books are obscure, low-demand manuals that might otherwise be pulped, and note that current copyright law perversely incentivizes private, destructive digitization while blocking open preservation. The thread raises broader concerns about long‑term knowledge loss, concentration of information inside corporate AI models, and calls for reforms such as shorter copyright terms, public digitization projects, or requirements to preserve and eventually release scanned texts.

What “rare books” means and how serious the loss is

  • Many note “rare” is loosely used: could mean low print runs, foreign-language titles, old manuals, or outdated technical books, not priceless first editions.
  • Skeptics argue the article withholds titles because they’re mundane (e.g., 1990s software guides), calling the outrage manufactured clickbait.
  • Others counter that even obscure or niche works (local histories, sermons, manuals, periodicals) can become historically valuable or uniquely important to specific research.

Destructive scanning, copyright, and legal backdrop

  • Several comments tie destruction directly to copyright strategy: digitizing and destroying purchased copies has been treated as fair use in at least one US case (Bartz v. Anthropic).
  • There’s dispute over what exactly the ruling permits; some say destruction is legally required, others say it just strengthens the fair‑use argument.
  • Strong criticism that copyright law now incentivizes private destructive digitization while punishing open digitization efforts like Google Books / Internet Archive.

Preservation vs. destruction and access

  • One side: AI labs are preserving content that would otherwise be landfilled; books are constantly weeded by libraries, dumped by estates, or pulped by publishers.
  • Other side: preservation “in model weights” is not preservation—no verbatim text, no provenance, and proprietary models can disappear or be altered.
  • Concern that multiple AI firms, all competing for the same dwindling pool, could destroy the last physical copies of some works, with no public record or guarantee of released scans.

Value of obscure print data for AI

  • Supporters: books generally contain higher‑quality, deeper writing than the web; frontier LLM work treats “more diverse data” as clearly beneficial.
  • Skeptics: marginal value of obscure out‑of‑print books is tiny compared to existing digital corpora; this looks more like competitive hoarding than necessity.

Libraries, markets, and scale

  • Many note massive ongoing book destruction predating AI, especially via library weeding and thrift‑store dumping.
  • Others reply that AI-driven scanning is qualitatively different: industrial scale, targeted at undigitized works, and done without public cataloging.

Corporate incentives and proposed remedies

  • Repeated worries about private “walled gardens” of knowledge and the power to quietly revise the digital record.
  • Suggested mitigations:
    • Commit to depositing scans into a public trust when copyright expires.
    • Publish lists of scanned/destroyed titles.
    • Government‑funded large‑scale digitization via national libraries.
    • Copyright reform (shorter terms, link protections to “in print” status, bulk licensing / clean rooms).

Reaction to the article and media framing

  • Some praise the investigation into the supply chain of bulk book buyers.
  • Others criticize 404 Media’s tone as rage‑bait and light on specifics (no titles, prices, or volumes), making the ethical stakes hard to judge.