Japan Goes All In: Copyright Doesn't Apply to AI Training

Japan’s move to exempt AI training from copyright restrictions is seen as a potential bid to become an AI powerhouse, but it raises sharp questions about the future of creative rights and business models built on intellectual property. Commenters distinguish between allowing models to learn from copyrighted data and holding developers or users liable when outputs closely reproduce protected works, citing cases like New York Times vs. OpenAI. Opinions range from viewing this as overdue copyright reform and a boost to innovation, to warnings that it could undermine artists, news organizations, and even privacy if broad data use becomes normalized.

Regulatory Arbitrage & Global Competition

  • Many expect “AI-friendly” copyright rules to create regulatory arbitrage if US/EU go the opposite way.
  • Some see Japan (and mentioned South Korea) as deliberately positioning to become AI hubs and avoid NYT-vs-OpenAI–style litigation.
  • Others warn that strict regimes will just entrench big players who can afford licenses or buy content sources.

Inputs vs Outputs: Where Copyright Bites

  • Broad agreement that the critical distinction is inputs vs outputs.
  • Supporters of Japan’s stance say: training on copyrighted material (even from “illegal sites”) is like human reading; infringement should be judged on output.
  • Critics respond that current models can and do emit near-verbatim text and images (NYT lawsuit, Getty watermark examples), so training can’t be treated as harmless in practice.

Impact on Creators and Business Models

  • Concerns that artists’ distinctive styles and news organizations’ expensive reporting are being commoditized and mechanized.
  • Some argue this is just “price of progress,” others insist society should be honest about the distributional harm.
  • Debate over whether style can or should be protected; one side says styles can’t be owned, the other sees this as hollowing out creator value.

Copyright Theory, Fair Use & Analogies

  • Recurrent analogies: LLMs vs humans, pencils, scanners, photocopiers, printing press; many participants criticize these as misleading, especially given speed and scale differences.
  • One camp says law should privilege human learning but not extend that privilege to machines; another says trying to draw that line is impractical and invites loopholes and Mechanical Turk workarounds.
  • Some want copyright radically shortened or abolished; others fear this effectively ends copyright altogether.

Privacy vs Copyright

  • A minority worry that permissive “training” rules could extend to sensitive data like health records.
  • Others counter that privacy is governed by separate laws and isn’t meaningfully affected by copyright carve-outs.

Uncertainty About the Japanese Policy

  • Late in the thread, some note the cited statement appears to come from older committee minutes, not clear, current binding policy, and the article is described as possibly misleading or over-editorialized.