When the infrastructure layer picks a side, the whole web shifts at once.
A Default Is a Decision
In mid-2025, Cloudflare announced that AI bot blocking would become the default setting for all new customers on its network. Publishers who had previously needed to navigate the dashboard and opt in to block crawlers like GPTBot, CCBot, and Google-Extended would no longer have to act at all — the block would arrive preconfigured. Existing customers were not automatically switched, but the company made the new default prominent in its product materials.


The practical reach of that decision is hard to overstate. Cloudflare sits in front of a substantial fraction of all web traffic, acting as reverse proxy, CDN, and DDoS shield for millions of domains. When Cloudflare changes a default, it changes the crawlable web in bulk — not through a standards process, not through legislation, but through a product toggle. No W3C working group, no IETF RFC, no robots.txt negotiation: a single company's dashboard setting, replicated across an enormous slice of the internet.
What the default blocks is a curated list of crawlers Cloudflare identifies as AI training bots. The list draws on user-agent strings and, where possible, verified crawler identities. GPTBot — OpenAI's crawler, disclosed in August 2023 — is included, as is Google-Extended, Google's opt-out handle for AI training data. CCBot, operated by Common Crawl, the nonprofit that maintains open web archives, is also on the list, a detail that drew pointed commentary from the open-data community, because Common Crawl's corpus underlies academic research as much as commercial model training.
From the record
The policy in brief
- Announced: mid-2025
- Scope: new Cloudflare customers, default-on; existing customers not automatically changed
- Crawlers blocked by default: GPTBot (OpenAI), Google-Extended, CCBot (Common Crawl) and others on Cloudflare's identified AI-bot list
- Mechanism: proxy/firewall block before the request reaches the origin server, distinct from robots.txt
Points of controversy
- Common Crawl's CCBot blocked alongside commercial training crawlers, despite Common Crawl's nonprofit, open-archive mission
- No standards body or legal process involved:a product default replicating at infrastructure scale
- Existing customers not changed automatically, but the nudge is now toward blocking by default
The robots.txt convention — a 1994 plain-text file carrying no legal enforcement — had been the primary tool publishers reached for when GPTBot appeared. Cloudflare's default represents a different layer of the same fight: not a polite note in a text file, but a firewall rule applied before a request reaches the origin server at all. Bots that disregard robots.txt directives still encounter Cloudflare's block at the proxy level; they simply receive a 403 or a CAPTCHA challenge and stop.
The policy shift reflects a broader argument Cloudflare made publicly: that AI training crawls impose real bandwidth costs on publishers who receive no compensation, and that acting at the infrastructure layer is both technically more reliable and more accessible to non-technical site owners than expecting everyone to edit a text file. Whether blocking Common Crawl indiscriminately costs the open web more than it protects any individual publisher remains a live dispute — one that the Data Provenance Initiative's measurements of rising blocking rates had already framed before Cloudflare moved the default.



