NY Times copyright suit wants OpenAI to delete all GPT instances
The New York Times’ lawsuit against OpenAI and Microsoft alleges that their AI models were trained on millions of copyrighted Times articles without permission, can sometimes reproduce them verbatim, and now compete with the paper’s subscription-based business. Commenters debate whether this training is protected “fair use” or closer to large‑scale piracy, drawing analogies to search engines, human learning, and past copyright cases, and arguing over the role of paywalls and terms of service. Many expect the case—or one like it—to set pivotal legal precedents for how generative AI may use copyrighted content, with high stakes for both big tech and smaller or open-source AI efforts.
Scope and goals of the NYT lawsuit
- NYT demands: destruction of GPT models trained on its content, deletion of training datasets, and a permanent injunction plus large monetary damages.
- Many see this as a maximalist opening move meant to force a licensing deal rather than actually “kill” GPT.
- Some think timing is tied to broader AI–publisher negotiations (e.g., with Apple and others).
Copyright, fair use, and scraping
- Central question: is using copyrighted articles for training itself infringement, or only infringing when outputs are too close to originals?
- Debate over fair use factors:
- Purpose/character: transformative learning vs. commercial substitution.
- Nature: news as partly factual but with creative expression.
- Amount: models ingest entire archives.
- Market effect: whether LLMs meaningfully replace NYT readership or licensing markets.
- Scraping legality and ToS: some cite precedent that public web scraping can be lawful; others note paywalls, changed NYT terms, and robots.txt as evidence of non-consent.
Memorization, regurgitation, and technical fixes
- NYT’s strongest examples are near-verbatim reproduction of paywalled articles and first paragraphs.
- One camp: this proves the model “contains” copyrighted works and that training is more like compressed copying than human-style learning.
- Other camp: these are rare overfitting edge cases that can be mitigated via lower temperatures, filters (n‑grams/bloom), RLHF, or pre‑summarization of sources.
- Additional concern: models hallucinate false NYT content, raising defamation and trademark-dilution angles.
Human vs machine analogies
- Supporters of OpenAI liken training to a person reading and later summarizing or being “influenced” by NYT, arguing architecture shouldn’t matter legally.
- Critics stress scale, reproducibility, and commercial deployment: millions of users, automated paywall bypass, and a single corporation monetizing others’ work.
- Repeated pushback against “if a human can do it, a model can” reasoning, noting different legal treatment for tools vs people.
Economic, competitive, and societal implications
- Many expect an eventual licensing regime: big publishers get paid directly, smaller creators maybe via pools; LLM access becomes more expensive.
- Concern that strong copyright rulings would advantage deep‑pocketed firms (OpenAI, MS, Apple, Google), hurt open-source and startups, and possibly push frontier work to jurisdictions with weaker IP enforcement.
- Others argue this constraint is necessary to preserve incentives for journalism and creative labor, and to prevent a few AI companies from privatizing the value of the web.
- Thread reflects broader discomfort with current copyright length and scope, but sharp disagreement on whether AI training should be carved out as clearly legal fair use.