NY times is asking that all LLMs trained on Times data be destroyed

The New York Times’ lawsuit against OpenAI and Microsoft, which demands that AI models trained on Times articles be deleted, is prompting a broader clash over how copyright applies to large language models. Commenters debate whether training on copyrighted text should be treated like human learning and fair use, or as large‑scale, uncompensated exploitation that harms creators whose work can be regurgitated verbatim. Many expect courts to focus on model outputs and market impact rather than training alone, with likely outcomes ranging from licensing and royalties to stronger limits that could entrench big tech firms and reshape how future AI systems are built.

Scope of the Lawsuit and “Destroy the Models” Demand

  • Thread centers on the NYT suit against OpenAI/Microsoft and a request that models trained on Times content be deleted.
  • Several comments argue courts can only order relief against named defendants, not “all LLMs,” so the prayer for broad destruction is mostly rhetorical leverage.
  • Others see the suit as a necessary pushback because individual creators cannot afford similar litigation.

Legal Status of Training on Copyrighted Material

  • Widely noted that legality is unsettled; multiple ongoing lawsuits are expected to clarify whether training is fair use.
  • One side: training = reading/learning; copyright restricts reproduction, not learning or remembering ideas. Analogies made to students, search engines, and ISP/browser caching.
  • Other side: LLMs are commercial tools, not humans; treating them like people is misleading. Using copyrighted text at scale for profit without permission is seen as infringement or theft.
  • Debate over whether model weights can themselves be “copies” of works if the model can emit near-verbatim text.

Outputs, Regurgitation, and Possible Remedies

  • Broad agreement that verbatim or near-verbatim reproduction of paywalled or copyrighted text is legally risky.
  • Proposed technical fixes: post-filters comparing outputs to NYT text; retraining without NYT data; or making models unable to reproduce long passages.
  • Some argue infringement should attach to specific outputs and their users, not mere existence of the model.

Ethical and Economic Perspectives

  • Strong split:
    • Some see LLMs as industrial-scale plagiarism, harming journalists, artists, and coders and enabling further wealth concentration.
    • Others see them as another technology layer (like VCRs, industrialization, or word processors) that builds on prior culture and ultimately boosts productivity.

Practical and Strategic Implications

  • Retraining without NYT content is described as extremely expensive, close to starting from scratch.
  • Licensing is seen by many as the likely endgame, but there’s concern it would set a precedent leading to many similar claims and favoring big incumbents with proprietary data.
  • Some suggest public-domain models or mandated licensing/royalty schemes; others propose, more radically, forcing models or even all published works to be available for training.