OpenAI and journalism
OpenAI’s response to The New York Times’ copyright lawsuit has reignited debate over whether using online text to train large language models counts as fair use, especially when models occasionally reproduce paywalled or copyrighted material verbatim. Commenters argue over key legal and ethical issues: if training on scraped data without consent is permissible, whether regurgitation can be dismissed as a “bug,” how much commercial impact AI has on journalism and other creators, and whether opt‑out policies or technical filters meaningfully address these concerns. Many see the case as a precedent-setting moment that could reshape both AI development practices and the economic incentives to produce original content on the internet.
Fair use and training data
- Major dispute over whether training on copyrighted text is fair use.
- Some argue it’s clearly transformative and analogous to prior cases (e.g., large-scale book scanning / search).
- Others say using people’s work to build a competing product is not fair use, especially when done on pirated or paywalled content.
- Debate over whether “publicly available” excludes paywalled or partially paywalled sources like major newspapers.
Regurgitation vs. transformation
- Many accept training might be fair use but see verbatim regurgitation as clear infringement, regardless of intent or rarity.
- Others argue that if a human can legally recall or paraphrase, an AI should be treated similarly; counterpoint is that mass, automated reproduction is fundamentally different.
- Question raised: if regurgitation is <1%, does that meaningfully weaken claims that models substitute for reading original sources?
Legal analogies and liability
- Comparisons made to:
- Reciting articles in public.
- Google/YouTube/Dropbox hosting user content under DMCA safe harbors.
- A “savant” reproducing books from memory vs. summaries/analyses.
- Some say inference output is not “user-generated” under DMCA and so platforms lack that shield.
- Debate about where responsibility lies: tool maker vs. downstream users and intermediaries.
Opt-out, transparency, and datasets
- Strong criticism of opt-out introduced only recently and not applied retroactively.
- View that “you can’t unlearn” is used as a self-serving excuse; others reply that full retraining is technically possible but expensive.
- Calls for models to at least disclose what datasets they were trained on, even if raw data can’t be released.
Technical mitigation
- Suggestions: output filters to detect copyrighted text, post-generation rewriting, building reference search against training corpus.
- Disagreement on feasibility: some say simple text filtering is doable; others point out LLMs lack explicit storage/indexes and can’t “unlearn” specific items reliably.
Impact on journalism and the internet
- Concern that if AI systems replace reading original sources, incentives to produce quality journalism and human-created content erode.
- Counterview: current regurgitation use case is marginal compared to existing paywall workarounds; chilling AI progress to appease publishers is seen as too high a cost by some.
Perception of OpenAI’s response
- Many see the blog post as PR aimed at the public and enterprise customers, not serious legal argument.
- Skepticism about claims that any single source (e.g., a major newspaper) is “not meaningful” to the model, given that all such sources are needed in aggregate.