Researchers put numbers to what publishers were doing by instinct

When publishers started blocking AI crawlers in 2023, the response looked chaotic — individual sites reaching for robots.txt with no coordination and no clear sense of whether it mattered. The Data Provenance Initiative, a consortium of researchers affiliated with MIT and other institutions, turned that chaos into a dataset.

A text editor open on a monitor showing a robots.txt file with GPTBot and CCBot disallow rules, office background
Code on a screen, parsing a feed. The instrument for refusing a crawler is plainer than this: two lines of text on the same server.Photo: Nemuel Sereti / Pexels

Their 2024 study examined robots.txt files across a large sample of web domains, tracking which sites had added AI-crawler disallow rules and how quickly those restrictions had spread since OpenAI published its GPTBot user-agent string in August 2023. The methodology was straightforward: crawl robots.txt files, parse the disallow directives by user-agent, and compare the state of the web before and after the GPTBot disclosure.

The numbers were striking. Within months of that August 2023 announcement, a measurable and accelerating share of high-quality web content had been placed behind explicit crawler restrictions. The research found that the domains blocking AI crawlers were not randomly distributed. News publishers and content sites with professional editorial operations moved fastest. Domains that Common Crawl — the nonprofit that has archived the web since 2008 and whose datasets underpin many large language model training runs — had rated as high-quality sources were disproportionately represented among the blockers.

From the record

What the research found

  • Study scope: robots.txt files across a large sample of domains, parsed by user-agent before and after August 2023
  • Key finding: high-quality domains,those rated well by Common Crawl, blocked AI crawlers at disproportionately high rates
  • Pattern: news and editorial sites moved fastest; the open residual web skewed toward lower-quality content
  • Spread: Google-Extended blocking tracked closely with GPTBot blocking, suggesting a generalised anti-AI-training stance
  • Limitation the researchers acknowledged: robots.txt compliance is voluntary; the study measured publisher intent, not crawler behaviour

Timeline

  1. August 2023OpenAI publishes GPTBot user-agent
  2. Weeks laterGoogle introduces Google-Extended
  3. 2024Data Provenance Initiative publishes findings documenting the acceleration of blocking across both agents

That finding carried a specific implication: the content most useful for training capable AI systems was precisely the content disappearing from the crawlable web first. The researchers described a dynamic in which quality and restriction were positively correlated, meaning the residual open web was skewing toward lower-quality material as the wall rose.

The Initiative also tracked the spread of blocking to Google-Extended, the user-agent Google introduced weeks after GPTBot, and found adoption curves that mirrored the OpenAI pattern. Publishers were treating the two agents similarly, suggesting the blocking behavior was a response to AI training generally rather than to any single company.

Small self-hosted server on a home-office shelf with ethernet cables routed neatly, indicator lights lit
Miniflux and FreshRSS will run on hardware of this order. That is the whole argument for hosting a reader yourself.Photo: panumas nikhomkhai / Pexels

What the Data Provenance Initiative did not measure was crawler compliance or effect. robots.txt carries no legal enforcement mechanism of its own, and whether any given AI company honored the directives was outside the study's scope. The research documented a signal — publishers pulling content away from AI crawlers at a measurable rate — without being able to confirm whether the signal was received.