Cloudflare's new AI traffic options for customers
Cloudflare’s new controls for AI-related traffic are raising questions about who gets to crawl the web, on what terms, and who ultimately pays for it. Commenters highlight that blocking “training” crawlers will soon also block major search bots like Googlebot because they share infrastructure, potentially cutting off vital search traffic while doing little to stop abusive or opaque scraping. Others worry that Cloudflare is positioning itself as a de facto gatekeeper and “tax collector” for access to web content, while alternative defenses like proof‑of‑work schemes remain imperfect and can degrade access for legitimate users.
Cloudflare’s new AI controls and defaults
- Robots.txt and similar signals are still seen as an “honor system”; many doubt AI crawlers will respect them.
- New Cloudflare options distinguish Training, Search, and Agent traffic; some like the added granularity, others think the only relevant distinction is “bot vs human.”
- For new domains, Training and Agent are blocked by default on ad pages; some worry about unintentionally losing AI traffic because many users never change defaults.
Impact on Googlebot and search traffic
- A key change: multi‑purpose crawlers (Googlebot, Applebot, Bingbot) are treated as AI trainers if any of their uses involve training, so blocking Training effectively blocks them.
- One user reports blocking AI training via Cloudflare halved their Google search traffic.
- Debate over whether Googlebot causes de facto DDoS: some say it’s conservative and most “Googlebot” traffic is fake; others report heavy load from verified Google IPs and little recourse.
Cloudflare’s role, incentives, and centralization
- Some see Cloudflare as “playing both sides”: selling tools for AI agents while also positioning itself as gatekeeper and future toll collector for pay‑per‑crawl.
- Concerns about a de facto centralized “tax collector” for web access, potentially favoring incumbents and raising barriers for new crawlers.
- Others argue Cloudflare is trying to chart a middle path, offering tools for permission and compensation rather than an all‑or‑nothing block.
Bot-blocking methods and user friction
- Alternatives like proof‑of‑work (e.g., Anubis) are debated:
- Proponents say PoW effectively rate‑limits large scrapers and preserves privacy.
- Critics say sophisticated bots with headless browsers and large compute budgets can solve PoW at scale, and some evidence is cited of AI crawlers adapting.
- PoW can severely degrade experience on low‑end devices; Cloudflare challenges can also block legitimate users, sometimes without a visible captcha.
Ethics and feasibility of controlling AI scraping
- Some want to block all bots except specific search/AI providers; others note Google is reducing outbound traffic via AI overviews, lowering incentives to allow crawling.
- There’s worry that blocking “AI” also blocks beneficial user‑controlled agents and accessibility tools.
- One view: technically, if content is readable at all, it can’t be meaningfully restricted from training; promises to separate “reading” from “training” are seen as misleading.
- Broader anxiety that unilateral moves by Cloudflare and large AI firms will reshape incentives, possibly pushing more of the web behind paywalls or closed gateways.