Things are about to get worse for generative AI

Generative AI systems like ChatGPT and DALL·E are increasingly able to reproduce copyrighted or trademarked characters and even news articles from vague prompts, raising the risk that users and AI providers unwittingly infringe IP law. Commenters debate who should bear legal responsibility—the model operators, the users, or the rights holders themselves—and whether existing copyright frameworks can or should be stretched to cover large-scale machine training on scraped data. Many expect a mix of outcomes: tighter licensing and filtering for commercial models, stronger enforcement by major media and entertainment companies, and a parallel ecosystem of “pirate” or open models trained on unrestricted or illicit datasets.

Liability and the “who gets sued” question

  • Strong disagreement over whether the primary infringer is:
    • The end user who publishes AI‑generated content.
    • The model provider who trained on copyrighted data and distributes outputs.
  • Analogies used:
    • Session musician whose riff copies “Stairway to Heaven.”
    • Contractor vs employer IP ownership.
    • Photoshop/Xerox as neutral tools vs a model that already “contains” protected works.
  • Some note vendors like OpenAI/Microsoft/AWS offer limited indemnity, implicitly accepting some risk.

Training on copyrighted data vs output infringement

  • One camp: training on copyrighted works is itself infringement; curators must ensure only properly licensed or public‑domain data.
  • Another camp: training is transformative “text and data mining” and likely fair use (or explicitly exempt in some jurisdictions), with infringement only at use/output.
  • EU and Japan cited as having specific TDM exceptions, though with opt‑out or “unreasonable prejudice” caveats; legal outcomes seen as unresolved/unclear.

Trademarked and copyrighted characters from vague prompts

  • Examples: “videogame plumber,” “golden droid from classic sci‑fi movie,” “animated sponge” producing clear Mario/C‑3PO/SpongeBob analogues.
  • Some say these prompts are effectively just indirect requests for specific IP, so it’s not surprising or exculpatory.
  • Others stress this makes it easy for users to unknowingly infringe; current models don’t warn or attribute.
  • Proposed mitigations:
    • Post‑processing filters or separate detection models (like YouTube ContentID).
    • Prompt‑ or output‑level blocking based on trademark databases or model‑based recognizers.
    • Critics argue this is technically brittle, legally murky, and prone to over‑blocking.

Impact on artists, incentives, and copyright policy

  • Many artists/creatives in the thread see GenAI as already harming freelance work and commissions, especially for individuals more than big rightsholders.
  • Counter‑view: people create for intrinsic reasons; past tech (photography, synths) didn’t kill art; creation will adapt.
  • Broader critiques:
    • Current copyright terms (life + decades) seen by several as excessive and favoring large corporations.
    • Proposals range from shortening terms and strengthening fair use, to abolishing copyright or shifting to patronage/gratuity models.
    • Others warn that gutting copyright would mostly empower big tech to monetize everyone’s work without compensation.

Regulation, geopolitics, and future trajectories

  • Some predict strong copyright‑based constraints will “kneecap” commercial GenAI in the US/EU while open/pirate models and non‑Western jurisdictions forge ahead.
  • Analogies drawn to Napster vs torrents: even if major services are constrained or forced to retrain, uncensored models and datasets will persist.
  • Speculation about outcomes:
    • Licensing deals and “AI taxes” on prompts or models.
    • Separate, stricter rules for high‑risk uses, with lighter treatment for private, noncommercial generation.
    • Possible chilling effect on trust in AI outputs if they are heavily post‑filtered and opaque.