Database of artists used to train Midjourney AI garners criticism

A leaked list of 16,000 artist names linked to Midjourney’s prompt system has reignited concerns about how generative AI models are trained on copyrighted artwork without consent. Commenters debate whether AI-generated images should qualify for copyright, how to distinguish “AI-assisted” from “AI-authored” work, and whether future regulation will meaningfully protect artists or simply entrench large platforms that can afford licensed datasets. Some see the shift as an inevitable technological upheaval akin to past automation, while others argue that models which can regurgitate training data cross a legal and ethical line that courts will eventually have to address.

Status of copyright for AI-generated works

  • Several comments note that, in the US, fully AI-generated images currently cannot be copyrighted; only the human-created parts of mixed works are protected.
  • Others stress this is nuanced: the Copyright Office allows registration of AI-assisted works where human contributions are sufficiently creative, but not works where AI is the primary author.
  • Debate over whether prompt-writing plus selection/iteration is enough authorship; some argue it’s like photography or Photoshop, others say a mere prompt isn’t sufficient.
  • Unclear thresholds: how much “touch‑up” or editing makes an AI-derived image copyrightable remains vague.

Enforcement, detection, and provenance

  • Many see a blanket “no AI art” or “AI art is public domain” rule as practically unenforceable because AI output is hard to detect, especially when edited.
  • Proposals include flipping the default (assume AI unless proven otherwise) and requiring creators to keep process artifacts; critics see this as burdensome and unrealistic for most artists.
  • Some suggest cryptographically authenticated editing histories for digital art; others point out these can be gamed or will be onerous across media types.

Training data legality and ethics

  • Strong disagreement on whether training on copyrighted works is lawful “fair use” or outright theft.
  • Some argue models are like compression; infringement occurs only when users force verbatim regurgitation. Others reply that memorization research shows models effectively store protected works, making the weights themselves infringing.
  • There’s debate over whether outputs are meaningfully “derivative,” how to attribute value to millions of training items, and whether artists should receive royalties.

Midjourney artist list specifics

  • Multiple commenters note the 16,000‑artist list is not the actual training set but a post‑hoc list (scraped from sources like Wikipedia) used to optimize style prompts.
  • Midjourney is described as trained broadly on “everything,” not targeted only at those artists, though that broad scraping is itself criticized.

Economic and policy implications

  • Some think the battle is already lost due to proliferating open models and permissive jurisdictions; others point to Napster/Spotify as evidence that law and licensing can reshape practices.
  • Concerns raised that stricter IP rules will favor large incumbents who can buy massive licensed datasets, further centralizing power.
  • Comparisons are drawn to historical automation (Luddites, mechanical looms); disagreement persists over whether this is a similar technological shift or uniquely exploitative because it required mass unconsented copying.