The New York Times is suing OpenAI and Microsoft for copyright infringement
The New York Times’ lawsuit against OpenAI and Microsoft over alleged copyright infringement by ChatGPT and Copilot has become a test case for how existing IP law applies to large language models. Commenters debate whether training on copyrighted text is more like human learning (and thus fair use) or like unlicensed copying at industrial scale, especially when models can reproduce paywalled articles or closely mimic a publication’s style. Many see high stakes for both sides: a ruling against OpenAI could force costly licensing and retraining, while a win for AI firms might accelerate pressure on already fragile news and creative business models.
Allegations and Evidence Discussed
- NYT lawsuit claims OpenAI/Microsoft copied “millions” of NYT articles, and that ChatGPT/Copilot:
- Can output NYT text verbatim.
- Closely summarize paywalled content.
- Mimic NYT’s expressive style.
- Commenters highlight complaint examples where prompts like “I’m paywalled out of [specific article], give me the first paragraph” allegedly returned near‑verbatim text, including Wirecutter reviews via “Browse with Bing,” stripped of links and referrals.
Training on Copyrighted Data vs. “Just Reading”
- One side: training is akin to reading/education; model weights are transformative statistics, not copies; using publicly reachable web text should be fair use.
- Other side: training is a commercial, industrial-scale reuse unlike human reading; inclusion without permission or payment is exploitation, especially when outputs substitute for the original.
Verbatim Regurgitation vs. Style and Summaries
- Broad agreement that:
- Simple “style imitation” and high‑level summaries are likely non‑infringing.
- Large verbatim or near‑verbatim passages, especially on demand and without attribution, are the NYT’s strongest claim.
- Dispute over how often this really happens, whether it’s a patched bug or intrinsic to large models, and how hard it is to trigger.
Search Engines vs. LLMs
- Some say search is symbiotic (sends traffic and ad revenue), while LLMs are parasitic (answer directly, no click‑through).
- Others note that Google already pushes into “answering on page” (snippets, Knowledge Graph) and has also faced publisher pushback.
Impact on Journalism and Incentives
- Concern: if LLMs free‑ride on reporting, they may undermine subscription/ads/affiliate models and reduce incentives for costly investigative work.
- Others argue journalism was already economically fragile; AI is another wave like Napster for music, and law should adapt rather than freeze AI.
Legal Uncertainty and Precedent
- Heavy discussion of US fair use’s four factors, especially “effect on the potential market.”
- Some expect courts to find training fair use but punish verbatim regurgitation; others think even training may be ruled infringing or require licensing.
- Widespread view that this and similar suits (vs. code copilots, book datasets, etc.) will set critical precedent.
Technical and Policy Proposals
- Ideas floated:
- Royalty/licensing schemes or “AI training marketplaces.”
- Taxing AI outputs instead of restricting inputs.
- Training only on public‑domain or opt‑in corpora.
- “Machine unlearning” to remove specific sources from trained models.
- Concerns that strict copyright enforcement could:
- Entrench big players who can afford licenses.
- Push cutting‑edge training to more permissive jurisdictions (e.g., Japan, China).