Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI
Unsealed court filings in the Authors Guild’s lawsuit against OpenAI and Microsoft allege the companies knowingly used pirated book datasets such as LibGen to train large language models, while internally worrying more about “optics” on sites like Hacker News and Twitter than legality. Commenters debate whether such large‑scale scraping and AI training is protected fair use or constitutes unprecedented copyright infringement and “theft of labor,” highlighting perceived double standards between how individuals and tech giants are treated under IP law. The thread also raises broader concerns about AI’s impact on creative livelihoods, the enclosure of the cultural commons, the flood of low‑quality “AI slop,” and the loss of human control and leverage in an economy increasingly mediated by powerful proprietary models.
Use of LibGen and “optics”
- Unsealed filings show staff discussing using LibGen’s large pirated book corpus and worrying mainly about bad “optics” if HN/Twitter noticed, not legality.
- Some see the “sketchy Russian website” quote as hypocritical given far larger-scale scraping/piracy by labs; others say the concern was PR (Russia, troll-farm associations), not copyright or LibGen’s ethics.
- Disagreement over LibGen’s framing: pirate site full of in‑copyright textbooks vs. de‑facto research library in poor countries.
Piracy vs. training and fair use
- Broad agreement that downloading books from LibGen/Anna’s Archive is copyright infringement.
- Major dispute: is training on lawfully acquired copyrighted works fair use?
- One side cites recent US cases (e.g., Anthropic) that found training on lawfully acquired books “transformative” and fair use, but not training on pirated copies.
- Others argue models can reproduce copyrighted text, so training+deployment should be treated as copying, not mere “reading.”
- Debate over whether current evidence shows routine verbatim regurgitation; some say it’s rare and blocked, others point to lawsuits (news, lyrics) and “slop” outputs.
Jobs, power, and capitalism
- One camp treats AI as another automation wave (cars vs horses, calculators vs human “computers”); job loss is normal disruption.
- Another argues AI is different: it appropriates human creative labor at scale, undermines workers’ leverage, and risks extreme concentration of power (post‑labor “feudalism”/“god‑kings”).
- Split between those who want to slow/stop AI vs. those who want to accept tech progress and focus on safety nets, retraining, or post‑work economics (UBI, “luxury space communism”).
Culture, attribution, and “slop”
- Strong concern that models erase authorship, flood markets with low‑quality AI books/images, and erode cultural signal, truthfulness, and craft.
- Others counter that AI output currently cannot match skilled long‑form writing and is mostly used for low‑end copy and genre filler.
- Some argue copyright should be shortened or abolished; opponents note it also underpins copyleft (e.g., GPL).
Inequality and enforcement
- Repeated comparison to harsh treatment of individual pirates vs. billion‑dollar labs paying manageable settlements; sense that law mainly protects capital.
- Some see big‑lab noncompliance as a lever to reform copyright; others fear it will simply entrench a corporate carve‑out.
Hacker News’ role & meta
- Staff explicitly worried about negative HN coverage, implying HN still matters for “optics” among tech elites.
- Thread also debates astroturfing, moderation, and whether AI companies actively try to shape HN opinion.