Judge Rejects Google's Attempt to DMCA Its Way Out of Being Scraped
A recent U.S. court ruling rejecting Google’s DMCA claim against SerpApi, a service that scrapes Google search results, is being seen as a key affirmation that public search result pages can legally be scraped. Commenters highlight the irony of Google, whose dominance was built on crawling the open web, trying to block others from indexing its own pages while offering no robust public search API. The ruling is viewed as having broader implications for AI training, competition in search, platform liability for scam ads, and the copyright status of indexes and model outputs.
Perceived Hypocrisy and Nature of Scraping
- Many see irony in Google, built on mass web scraping, trying to block scraping of its own SERPs.
- Some argue Google’s crawling is “gentler” (robots.txt, no CAPTCHA breaking) versus SerpAPI’s alleged use of distributed proxies and scraping of bot-blocked sites.
- Others note Google’s market power means sites are effectively forced to allow its bots, so the “respect for robots.txt” is only part of the story.
Legal Status of Scraping Search Results
- Several commenters say scraping Google SERPs has likely always been legal, though Google will still fight it technically and contractually.
- A key takeaway: DMCA is about copyrighted content; bare search indexes/URLs may not qualify, so DMCA is a weak tool here.
- Some point out U.S. and EU antitrust moves already pushing Google to share search data with “qualified competitors.”
Search APIs, Access, and Competition
- Frustration that Google killed its public search APIs, pushing developers toward scrapers like SerpAPI.
- Some wish governments mandated public search indexes.
- Gemini “search grounding” is seen as too restricted and tightly controlled.
Implications for AI Training and Models
- Discussion that this precedent might imply AI model weights and distilled outputs are also hard to copyright, but this is flagged as speculative/unclear.
- Questions raised about how “frontier labs” have historically crawled data for training and how transparent they are.
Ads, Scams, and Platform Liability
- Strong concern about scammy ads in SERPs (e.g., visa/registry scams).
- Debate over whether platforms should be liable for fraud in their ad networks:
- One side wants strong liability akin to DMCA-style safe harbors with real enforcement.
- Others argue primary blame should remain on scammers, not platforms as “private police.”
- Section 230 is contested: some say it’s misapplied, others say it’s structurally broken for ad-driven platforms.
Copyright, DMCA, and Database Rights
- DMCA widely criticized; some even call for abolishing copyright, others warn of economic shock.
- EU database rights vs. U.S. originality standard are contrasted; grey area over whether ranked search results are “just facts” or creative selections.
- Maps and PageRank are discussed as edge cases of what’s copyrightable.
Google’s Moat and Future of Search
- Split views:
- Some insist Google’s moat (default status, non-blocked crawler, YouTube, cloud) remains very strong.
- Others believe LLM-based assistants are rapidly replacing classic search for many queries.
- A subset claims almost everyone they know has moved off Google; others report nearly everyone they know still uses it.
- Speculation that this ruling is partly about limiting how OpenAI/Anthropic access live Google results via intermediaries like SerpAPI, though details remain unclear.