The Announcement That Changed the robots.txt Graph

On 7 August 2023, OpenAI published documentation identifying GPTBot as the user-agent string its crawler used to harvest training data from the public web. The disclosure was straightforward — a named agent, a stated purpose, and an acknowledgement that publishers could block it via robots.txt. What followed was less straightforward: one of the fastest coordinated blocking responses the open web had ever produced.

Within weeks of that announcement, site operators were adding User-agent: GPTBot / Disallow: / to their robots.txt files in numbers that analysts described as unprecedented for a single crawler. Google followed OpenAI's lead in September 2023, naming Google-Extended as the user-agent for its AI training and data-improvement systems, distinct from the Googlebot that handles search indexing. A new disallow pattern appeared on millions of sites almost immediately afterward.

A text editor open on a monitor showing a robots.txt file with GPTBot and CCBot disallow rules, office background
Code on a screen, parsing a feed. The instrument for refusing a crawler is plainer than this: two lines of text on the same server.Photo: Nemuel Sereti / Pexels

What the Data Actually Showed

The Data Provenance Initiative, a research project tracking the composition and accessibility of machine-learning training datasets, measured the rate of blocking across prominent web domains and found that by mid-2024 a substantial and growing proportion of high-quality web content — the kind that feeds into curated datasets like C4 and The Pile — had been placed behind robots.txt restrictions aimed specifically at AI crawlers. The project's core finding was that the domains producing the most linguistically rich, carefully edited text were blocking at higher rates than average, meaning the signal AI systems most needed was becoming the hardest to reach.

Cloudflare's crawl data told a complementary story. The company, which sits between the public internet and a significant portion of its infrastructure, reported in 2024 that AI crawlers had become among the most frequently blocked categories of traffic it observed — and that GPTBot specifically ranked high on the list of agents publishers were naming in firewall and robots.txt rules. That observation fed directly into Cloudflare's 2025 decision to make AI bot blocking a default setting for new customers, a policy shift that moved the question from opt-in to opt-out.

From the record

The timeline

  1. 7 August 2023OpenAI publishes GPTBot documentation and user-agent string
  2. September 2023Google announces Google-Extended for AI training, separate from Googlebot
  3. Mid-2024Data Provenance Initiative publishes blocking measurement findings
  4. 2025Cloudflare makes AI bot blocking a default for new customers

What robots.txt can and cannot do

  • Can blocknamed crawlers whose operators have committed to honouring the convention
  • Cannot blockcrawlers operated by parties who ignore the file
  • Cannot undotraining data already collected before the block was placed
  • Cannot verifywhether a compliant operator is actually respecting the directive

The Oldest Tool in the Stack

robots.txt dates to 1994 — a plain-text convention developed before the commercial web existed, carrying no legal enforcement mechanism and relying entirely on crawler compliance. Its role in the AI training dispute exposed both its reach and its limits simultaneously. On the reach side: the convention is universally understood, costs nothing to deploy, and works against any crawler whose operator has committed to honouring it. OpenAI stated it would respect GPTBot exclusions. Google made the same commitment for Google-Extended. That voluntary compliance is precisely why the convention worked at all.

On the limits side: robots.txt blocks compliant crawlers only. It offers no redress against operators who ignore the file, no visibility into whether compliance is actually occurring, and no record of what was collected before the block was placed. Publishers who added GPTBot directives in August 2023 had no way of knowing how many times their content had already been visited, parsed, and absorbed into a training corpus. The directive was a lock fitted after the door had been open for years.

Small self-hosted server on a home-office shelf with ethernet cables routed neatly, indicator lights lit
Miniflux and FreshRSS will run on hardware of this order. That is the whole argument for hosting a reader yourself.Photo: panumas nikhomkhai / Pexels

What the 2023 blocking wave clarified, more than anything, was the structural asymmetry at the heart of the arrangement. Publishers had one lever — a thirty-year-old text file — and the training runs had already happened. The wave of disallow directives was real, measurable, and widely adopted. Whether it arrived in time to matter is a different question, and one that robots.txt was never designed to answer.