Google cut a deal with Reddit for AI training data

Google’s $60M-per-year deal to license Reddit data for AI training is prompting debate over who should benefit when user-generated content is monetized. Commenters question whether Reddit’s trove of posts is actually high‑quality or heavily astroturfed, and whether this move will further erode the “open internet” by pushing more conversation into private or paywalled spaces. Others note that Reddit’s terms of service already allow broad reuse of content, frame the deal as IPO-driven monetization similar to ad sales, and worry that contributors will see none of the financial upside.

Deal Size, Value, and Exclusivity

  • Many see $60M/year as cheap given Reddit’s centrality to search and its planned $5B IPO valuation; some estimate Reddit queries are a large share of their own Google use.
  • Others argue you can’t compare annual license fees to total company value; Reddit can resell data to many AI vendors.
  • Several suggest the payment is for incremental API access and legal cover, since Google already indexes Reddit; exclusivity is widely doubted.
  • A minority view: paying for data that users created for free is “obscene,” especially without user revenue sharing, while others reply that hosting and community-building justify Reddit’s cut.

Data Ownership, Compensation, and ToS

  • Multiple comments highlight Reddit’s ToS granting it broad, perpetual, sublicensable rights; users effectively consent to data resale.
  • Some argue if people wanted to be paid, they wouldn’t post publicly for free; others say they posted for discussion and ad-supported hosting, not as training data.
  • Concern that this shifts the internet from “marketplace of ideas” to raw material for AI and shareholders, eroding goodwill.
  • Comparisons to StackOverflow and advertising: debate over whether data-for-ads was a “fairer” implicit bargain than data-for-AI.

Impact on the Open Web, Privacy, and Platforms

  • Several think AI scraping of any public content is now inevitable; expect more discussion to move to private or semi-private spaces (Discord, etc.), though those could also monetize data.
  • Others see this centralization of both data and compute at a few firms as dangerous and under-regulated.
  • Some frame the move as IPO prep and long-term API monetization, including locking down public APIs and search access.

Quality and Bias of Reddit Data

  • Strong skepticism that Reddit is good training data: heavy astroturfing, bots gaming votes, echo chambers, and “confidently wrong” hive-mind answers.
  • Concerns about political and ideological skew, plus aggressive moderation agendas in some subreddits, potentially biasing AI.
  • Jokes that this will make AI more sarcastic, moralistic, and wrong, but also recognition that many LLMs already used older Reddit dumps.

User Responses and Countermeasures

  • Some say they’ll reduce or stop contributing to Reddit, “voting with time.”
  • Tools for mass-deleting or scrambling past Reddit content are shared as a defensive move, though doubts remain about what Reddit retains and resells.