Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

A US judge has approved a $1.5B settlement requiring Anthropic to pay about $3,000 per title for using pirated ebooks to train its Claude AI models, while affirming that training on lawfully acquired books qualifies as fair use. Commenters are sharply divided over whether this is a meaningful remedy or merely a “cost of doing business” that entrenches big AI labs by pricing smaller players out of access to high‑quality training data. The conversation broadens into critiques of modern copyright enforcement, the role of publishers versus authors in capturing compensation, and whether current law is adequate to protect human creators as AI systems learn from and potentially regurgitate their work.

Scope of the ruling

  • Thread repeatedly stresses: the court approved a settlement for piracy, not for “training on books.”
  • Training LLMs on lawfully acquired books was held to be fair use; the infringement was keeping a large, central library of pirated ebooks (e.g., from LibGen).
  • Separately, Anthropic also bought millions of physical books, destructively scanned them, and discarded the originals; this was ruled fair use because the copies stayed internal and the print copies were paid for and destroyed.

Size and meaning of the settlement

  • $1.5B (~$3,000 per book) is seen by some as a slap on the wrist and pure “cost of doing business” for a near‑trillion‑dollar company.
  • Others argue it’s ~100× the cost of buying the books and thus a reasonable civil damages figure, especially since it avoids the risk of a trial.
  • Typical split mentioned: about 50/50 between author and publisher per eligible title; lawyers get a sizable but reduced share; class reps get small bonuses.
  • Many note this is a voluntary civil settlement, not a criminal case; no jail was ever on the table here.

Fair use, memorization, and derivative works

  • One camp: training is analogous to humans learning from books; LLMs generally don’t output full works verbatim, so this is “transformative” fair use.
  • Other camp: models can regurgitate long copyrighted passages or near‑verbatim books in some cases, so calling this “transformative” is dubious.
  • Debate over whether profitability from models should trigger royalties or profit‑sharing, and whether outputs are “derivative works.”

Distribution of benefits and harms

  • Authors are generally poorly paid; many books never earn out their advance. Some say $3k/title is more than many authors ever see; others call it token money versus Anthropic’s gains.
  • Concern that most money flows to publishers and lawyers, not individual writers.
  • Several argue this outcome entrenches big labs: only capital‑rich players can afford massive book purchases and legal risk, raising barriers for open and small‑scale models.

Broader copyright and policy debates

  • Strong disagreement over copyright’s value: some want it shortened or abolished; others see it as essential to fund creative work.
  • Some propose ongoing royalties or even public‑utility style funding for training data; others insist on open‑weight releases if copyrighted works are used.
  • Anger over perceived double standards: individuals once ruined for piracy vs corporations now paying manageable settlements while keeping their models.