Judge rejects most ChatGPT copyright claims from book authors

A US judge has thrown out most copyright claims by book authors against OpenAI, leaving primarily a California “unfair competition” claim over using copyrighted works for training without permission. Commenters debate whether training large language models on pirated or unlicensed texts should count as infringement, contrasting long-standing arguments against media piracy with the current leniency shown to big AI companies. The exchange highlights unresolved questions about how existing copyright and fair use principles apply to AI training, the asymmetry between human learning and machine learning in law, and the potential impact on both open-source AI models and working authors.

Scope of the Ruling and Main Surviving Claim

  • Many see the key surviving issue as the state “unfair competition” claim about using copyrighted works for training without permission.
  • Several argue the article title is misleading because, even with most copyright claims dismissed, this remaining claim could be highly consequential.
  • Some note the judge was reluctant to treat all LLM outputs as infringing without specific evidence tying outputs to particular works.

Piracy, Copying, and Fair Use

  • Strong debate over whether using pirated copies for training should be clearly illegal, given decades of anti-piracy rhetoric.
  • One side: downloading/using pirated copies is effectively tolerated; only distribution is really enforced.
  • Others counter that unauthorized reproduction (including downloading) is infringement, even if rarely prosecuted, and that civil vs criminal enforcement is being conflated.
  • Some argue that if training is fair use, the legality of how the copy was obtained may not matter.

Human Learning vs LLM Training

  • Repeated analogies to humans reading books: if people can learn from books and later create, why can’t models?
  • Counterpoint: law treats human memory as “intangible” and not a “copy,” whereas computer storage is a fixed, tangible medium.
  • Several insist human cognition is legally special; machines don’t automatically inherit those exceptions or rights.

Ideas vs Expression and Derivative Works

  • Broad agreement that copyright protects expression, not ideas, procedures, or abstract concepts (citing U.S. statutory language).
  • Example discussion: using ideas from science fiction to build real technology is fine; reproducing the text is not.
  • Multiple comments stress that LLM outputs should be assessed for actual copying/market substitution, not just the fact that training ingested copyrighted works.

Economic and Competitive Implications

  • One camp fears requiring licenses for training would cement dominance of large firms that can pay for massive corpora; open models would suffer.
  • Another camp argues “just pay for it” (books, news, images) is fair and would apply equally to all actors, so not inherently anti-competitive.
  • Some expect, if licenses are required, a Spotify-like regime where media companies earn billions, creators get little, and only a few AI vendors survive.

Perceived Asymmetry and Legitimacy of Copyright

  • Frustration that harsh anti-piracy arguments used against individuals now seem softened when applied to large AI companies.
  • Some think if courts bless large-scale training on unlicensed works, it will de facto weaken copyright’s legitimacy for future piracy cases.

Regulation, Future Law, and Policy Direction

  • Several commenters think current copyright frameworks are ill-suited to LLMs and that new legislation is inevitable.
  • Disagreement over whether special “AI training rights” or explicit training exceptions should be created.
  • Some predict that powerful media interests will not allow “free exploitation” of all copyrighted works; others hope for narrow AI-specific carve-outs.

Quality and Source of Training Data

  • A few argue that books are higher-quality training data than general web text and note that many humans also rely on pirated ebooks.
  • Others emphasize that if books are used, AI companies should obtain them legally and/or license them, rather than scrape pirated copies.