Microsoft exec called AI scraping 'the largest theft of labor in human history'

An unsealed court filing revealing a Microsoft executive calling AI training “the largest theft of labor in human history” has ignited debate over whether large language models built on scraped web content constitute theft, copyright infringement, or transformative fair use. Commenters clash over analogies to human learning, piracy, and historical slavery, the hypocrisy of tech firms that once fought piracy now mass-ingesting copyrighted works, and the asymmetry between individuals punished for infringement and corporations rewarded for it. Many argue that even if training is legal, it devalues creative work, concentrates power in a few AI vendors, and may justify new regimes of taxation, royalties, or open-weight mandates to rebalance benefits toward the public.

Is AI training “theft” or just copyright infringement?

  • Many argue scraping for training is “the largest theft of labor”: models extract value from millions of works without consent or pay, then compete with the creators.
  • Others insist it’s not “theft” in the strict sense: originals remain, copying is copyright infringement at most, and law traditionally protects copies, not “ideas/facts/patterns.”
  • A recurring counterpoint: human learning from books is legal; opponents respond that a non‑rival human mind is not comparable to a massively scalable machine learner used commercially.
  • Legal status is described as unsettled: some courts have found aspects of training to be fair use; fair use vs market harm remains central and contested.

Scale, automation, and double standards

  • Commenters stress scale changes everything: one person pirating vs corporations hoovering the web, paywalled and even pirated datasets, then monetizing outputs.
  • Many note hypocrisy: individuals have been heavily punished for piracy; AI labs doing similar acts at massive scale are rewarded with trillion‑dollar valuations.
  • Labs’ complaints about model “distillation theft” are viewed as especially hypocritical, since their own models are trained on unlicensed material.

Impact on creators, culture, and labor

  • Strong concern that AI undercuts the market for writers, artists, coders, researchers, musicians, etc., by selling cheap substitutes built from their work.
  • Some open‑source and blog authors say they regret publishing publicly; they now see their unpaid work powering tools that may end their careers.
  • Others argue aggregation and remixing is how culture has always progressed; copyright maximalism itself is framed as a historical aberration driven by large media firms.
  • Fear that original web content will disappear as sites die (costs, bot load, loss of traffic), with knowledge effectively “trapped” in proprietary models and buried under AI-generated “slop.”

Commons vs property and redistribution

  • One camp welcomes AI as the ultimate “information wants to be free” moment and a partial unwinding of overextended IP, especially for works that arguably should be in the public domain.
  • Another camp insists IP is a legitimate right; if training on copyrighted work is allowed, then models and weights should also be treated as part of the commons.
  • Proposed remedies include: heavy taxes or profit‑sharing on AI firms, creator royalty schemes tied to training data, or universal benefits funded by AI gains.

Regulation and structural proposals

  • Ideas range from:
    • Require consent and licensing for training, with opt‑outs that actually work.
    • Mandate disclosure of training data and release of older model weights.
    • Declare model outputs uncopyrightable and potentially subject to original authors’ claims.
    • Cap company size or compute, or even ban certain proprietary frontier models.
  • Some note enforcement will be politically hard, especially given geopolitical competition and corporate lobbying.

Open vs closed models and distillation

  • Open‑weight models are cited as partial democratization, but critics say open weights trained on unlicensed data are still “laundered theft.”
  • Several defend distillation as a “human right” to work with models trained on humanity’s corpus; others worry it further entrenches the initial uncompensated appropriation.

Broader societal concerns

  • Threads explore:
    • Concentration of wealth and power if AI automates both mental and physical labor.
    • Environmental and infrastructure costs of huge data centers.
    • Analogies to surveillance: “public” data use becomes problematic when automated and centralized at extreme scale.