The convention that runs on trust, not teeth

In 1994, a Dutch developer named Martijn Koster posted a draft to the www-talk mailing list proposing a simple convention: place a plain-text file at the root of a web server, and crawlers would consult it before indexing the site's contents. No protocol layer enforced this. No certificate was required. The file would work, if it worked at all, because the parties reading it agreed to behave. Thirty years on, robots.txt is still a courtesy note — technically universal, legally close to nothing.

The specification describes a robots.txt file as a set of directives matching user-agent strings to allowed or disallowed paths. A Disallow: / blocks all crawlers from all paths, in theory. In practice, it blocks compliant crawlers. The distinction matters more with every passing year, because the set of crawlers that respect robots.txt and the set of crawlers scraping the web for AI training data are not the same set.

A text editor open on a monitor showing a robots.txt file with GPTBot and CCBot disallow rules, office background
Code on a screen, parsing a feed. The instrument for refusing a crawler is plainer than this: two lines of text on the same server.Photo: Nemuel Sereti / Pexels

What hiQ v. LinkedIn decided — and what it did not

The most-cited legal case in web-scraping discussions is hiQ Labs v. LinkedIn, litigated through the Ninth Circuit for the better part of a decade. hiQ, a workforce analytics company, scraped publicly visible LinkedIn profile data. LinkedIn sent cease-and-desist letters and deployed technical countermeasures. hiQ sued for injunctive relief, and LinkedIn counterclaimed under the Computer Fraud and Abuse Act, arguing that scraping constituted unauthorised access to a protected computer.

The Ninth Circuit's 2019 ruling found that the CFAA's definition of "unauthorised access" likely does not extend to scraping publicly accessible data — information that anyone with a browser can see without logging in. The Supreme Court vacated and remanded in light of its 2021 Van Buren ruling, and the Ninth Circuit reaffirmed its core reasoning in 2022. What the case established is narrow: a strong argument that the CFAA is probably the wrong instrument to prohibit scraping of public pages. What it did not establish: a general right to scrape, a ruling that Terms of Service are unenforceable, or any guidance on how copyright or contract law applies to the downstream use of scraped material for model training.

From the record

Timeline of key events

  1. 1994Martijn Koster proposes the robots.txt convention on the www-talk mailing list
  2. 2019Ninth Circuit rules in hiQ v. LinkedIn that the CFAA likely does not cover scraping of public pages
  3. 2021Supreme Court remands in light of Van Buren; 2022 Ninth Circuit reaffirms core reasoning
  4. August 2023OpenAI discloses GPTBot user-agent string; Google-Extended follows weeks later
  5. 2023Data Provenance Initiative measures rising AI-crawler blocking via robots.txt
  6. 2025Cloudflare makes AI bot blocking a default for new customers

The enforcement gap

  • robots.txtadvisory only; no protocol enforcement; relies on crawler compliance
  • Terms of Servicecontractual, user-accepted; more legally durable, but contested
  • CFAAComputer Fraud and Abuse Act; Ninth Circuit finds it likely does not reach public-page scraping
  • Network-layer blockingthe only mechanism that actually stops a non-compliant crawler

robots.txt featured in the case only in the background. It is a convention, not a contract. A site's Terms of Service — a document users accept — carries far more legal weight, and even that weight is contested jurisdiction by jurisdiction.

The compliance problem is getting worse

For most of the web's history, the major crawlers — Googlebot, Bingbot, the Internet Archive's crawler — treated robots.txt compliance as a matter of professional legitimacy. Violating it would damage relationships with publishers and, in Google's case, the entire ecosystem its business depends on. That incentive structure held the convention together.

Small self-hosted server on a home-office shelf with ethernet cables routed neatly, indicator lights lit
Miniflux and FreshRSS will run on hardware of this order. That is the whole argument for hosting a reader yourself.Photo: panumas nikhomkhai / Pexels

The AI training wave has introduced a different kind of crawler. Some operators of large language model training pipelines have released named user-agent strings so publishers can block them: OpenAI disclosed GPTBot in August 2023, and Google followed with Google-Extended weeks later. Both respect robots.txt — when addressed by name. The Data Provenance Initiative, a research coalition, measured in 2023 that a substantial and growing share of high-quality web content had been blocked against AI crawlers by the end of that year, using exactly this mechanism.

But a crawler that does not announce itself, or that rotates through generic user-agent strings, is invisible to any robots.txt rule that depends on a named match. The note is still polite. The recipient has simply stopped reading it.

For most of the web's history, the major crawlers — Googlebot, Bingbot, the Internet Archive's crawler — treated robots.txt compliance as a matter of professional legitimacy.

Cloudflare moved to make AI bot blocking a default for new customers in 2025, which is a technical response to the compliance gap: since robots.txt cannot be enforced, the intervention shifts from the advisory layer to the network layer, where actual gatekeeping is possible.

The situation reveals what robots.txt always was — a social contract among parties who share an interest in its maintenance. When that interest diverges, the file is just text.