'Impossible' to create AI tools like ChatGPT without copyrighted material, OpenA
Debate over OpenAI’s claim that tools like ChatGPT can’t be built without copyrighted material centers on whether training AI on such data is more like human learning or like copying into a database that can reproduce protected works. Commenters clash over legal and moral responsibility: some argue training should be fair use as long as outputs aren’t verbatim, while others see large-scale, commercial AI models as exploiting creators’ work without consent or payment. The exchange highlights broader worries about copyright duration, the fate of open-source models, corporate power, and how future laws should balance innovation with intellectual property rights.
AI vs Human Learning and Rights
- Big dispute over whether LLMs “learn like humans.”
- One side: both humans and LLMs learn from large volumes of language, so legally/ethically they should be treated similarly when consuming copyrighted works.
- Other side: implementation is fundamentally different; “consuming” text in a neural net is copying into a database-like system, unlike a human brain.
- Broader philosophical split:
- Some argue moral rights derive from consciousness/self‑awareness, not being biologically human, so advanced AIs should eventually get similar treatment.
- Others insist law is made by and for humans; machines are property, not rights‑holders, and expanding rights to AIs wastes political capital and masks corporate interests.
Copyright, Training Data, and Fair Use
- One camp: training on copyrighted data is akin to reading; non-infringing as long as the system doesn’t regurgitate works verbatim.
- Opponents: loading copyrighted material into training corpora is copying and adaptation, directly within copyright’s scope, regardless of the label “training” or “learning.”
- Some suggest content should be licensed; others worry this cements a monopoly for rich firms and crushes open‑source efforts.
Output, Memorization, and Infringement
- NYT examples of GPT reproducing articles are cited as proof of memorization and infringement.
- Counterpoint: many examples rely on prompts feeding part of the article; critics call this “cheating,” analogous to baiting a tool into infringement.
- Ongoing question: is liability on the model provider, the user who prompts it, or both?
Business Models, Scale, and Profit
- Several argue that what’s tolerated for individuals (piracy for personal use, private models) becomes unacceptable when monetized at massive scale.
- Some defend OpenAI’s “move fast, ask forgiveness later” approach as necessary to show what’s possible; others say they clearly exploited others’ work for profit and must now compensate.
Regulation, Jurisdictions, and Future Impact
- Concern that strict licensing requirements could:
- Undermine open models,
- Lock in Silicon Valley incumbents,
- Or handicap future embodied, continually learning AIs.
- Japan’s permissive stance on AI training is noted as a contrasting precedent.
- Underlying theme: current copyright terms may be too long, starving the public domain needed for modern AI and other public‑good uses.