An update on Wayback Machine access

Waves of high-volume scraping and bot traffic are overwhelming the Internet Archive’s Wayback Machine, leading to aggressive rate limiting that now blocks many ordinary users alongside abusers. Commenters weigh potential responses — from paid or metered access, stronger bot-detection tools, and login requirements to new standards for identifying and governing automated crawlers — while warning that copyright risk, AI training demand, and a “tragedy of the commons” dynamic could push more of the web, and even the archive itself, behind walls.

Traffic Surge and Scraper Abuse

  • Many think the Wayback Machine is being hammered by scrapers using it to bypass paywalls, bot blocks, and captchas, including via commercial scraping APIs that advertise “Wayback fallback.”
  • This extra load leads some sites to fully opt out of archiving to avoid indirect scraping.
  • Others note IA’s coverage is often partial, so it’s unclear how useful it really is for bulk scraping.

Paid Access, Copyright, and Micropayments

  • Some propose a paid, high-volume endpoint or bulk-download service to fund capacity.
  • Pushback: charging could weaken IA’s fair-use position by making use “commercial” and anger rightsholders, especially after IA’s recent book-lending lawsuit.
  • Debate over whether charging “for bandwidth, not content” is legally safer is unresolved and flagged as a “lawyer question.”
  • Micropayments and “downloader pays” models are discussed; critics point to user hatred of micropayments, regulatory overhead, and prior failures.

Bots vs. Humans; Blocking Side Effects

  • Many users report frequent 429 errors, especially from corporate networks, VPNs, IPv6 ranges, and some browsers, making normal use painful.
  • IA’s request that blocked users email OS/browser/IP is seen by some as reasonable signal-gathering, by others as impractical and user-hostile.
  • Several operators of public sites describe similar struggles: scrapers via residential proxies dominate traffic, making IP-based blocking and user-agent checks unreliable.

Agents, robots.txt, and “Good” vs “Bad” Scraping

  • Strong disagreement over whether all scraping is morally equivalent. Some say ends are the same; others argue that archiving/search scraping clearly benefits the public, while high-volume commercial scraping is parasitic.
  • Extensive debate on whether robots.txt should constrain LLM agents or only traditional crawlers; consensus is unclear.
  • AI agents that fetch many pages quickly blur the line between “human-like” browsing and industrial scraping.

Alternative Archives and Trust

  • Archive.org and archive.today are contrasted: one “cooperative” and robots-respecting vs. one “guerrilla,” paywall-bypassing, and accused of editing content and weaponizing users’ browsers in a DDoS—claims supported by linked Wikipedia/Ars Technica material.
  • Some still prefer archive.today for reliability or paywall circumvention despite trust concerns.

Values, Future, and Support

  • Many express strong support for IA as a unique public-good institution preserving internet history and donate or pledge to donate.
  • Others fear a “tragedy of the commons”: rising abuse may force logins, paywalls, or heavier gatekeeping, undermining open access.
  • Proposed longer-term directions include better bot standards (robots.txt v2, Web Bot Auth), decentralization/mirrors, and possible regulation of abusive scraping.