A License for the Thing robots.txt Was Never Designed to Cover
The robots.txt convention, written in 1994, tells crawlers where not to go. It says nothing about what a crawler may do with what it finds, how many times it may return, or whether commercial use requires any kind of agreement. That gap has widened into a chasm as AI training datasets have absorbed the public web, and Really Simple Licensing — RSL — is one attempt to fill it.
RSL is a proposed machine-readable licensing framework designed to be embedded in a site's HTTP headers or a dedicated endpoint, stating terms that crawlers would be expected to honour: permitted uses (research, search indexing, AI training), rate limits, attribution requirements, and whether commercial exploitation requires a separate agreement. Where robots.txt speaks only in allowed/disallowed binary, RSL attempts a richer vocabulary — closer to how Creative Commons licenses distinguish between share-alike, commercial, and derivative uses, but aimed at the crawl itself rather than at human readers or downstream publishers.


The specification draws on a straightforward premise: that publishers have legitimate interests the current toolchain cannot express, and that AI companies harvesting training data at scale are extracting value without a clear legal or contractual relationship. The Data Provenance Initiative's 2024 measurement work documented the rapid rise in web blocking following GPTBot's disclosure in 2023, showing that publishers are reaching for blunt instruments — total denial — because fine-grained instruments do not exist. RSL is a bid to create them.
The practical problem is enforcement. Creative Commons licenses work because copyright law backs them; a downstream user who violates the terms faces legal exposure. RSL, in its current form, is a convention, not a legal instrument. A crawler that ignores a robots.txt file also ignores an RSL declaration. Nothing in the HTTP stack compels compliance, and small publishers have no litigation budget to test the theory. The framework's proponents argue that major AI developers — facing regulatory scrutiny in the European Union and growing litigation in US courts over training data — have commercial reasons to adopt a voluntary standard that demonstrates good-faith effort. Whether that argument holds in negotiation is unproven.
From the record
Timeline of the gap RSL is trying to close
- 1994robots.txt introduced as a crawler courtesy convention
- 2023OpenAI discloses GPTBot; publisher blocking accelerates
- 2024Data Provenance Initiative publishes quantitative measurement of web blocking rates
- 2024–25RSL proposed as machine-readable licensing layer for crawl terms
The enforcement problem in brief
- robots.txtconvention only; no legal force
- Creative Commonsbacked by copyright law; violations are actionable
- RSL (current form)convention only; compliance is voluntary; no litigation pathway for most publishers
There is also the interoperability question. The web's content-licensing landscape already includes robots.txt, the emerging ai.txt convention, Cloudflare's bot controls, and platform-specific terms. A publisher managing four separate signals for four categories of crawler is already in complexity territory that strains smaller operations. RSL adds a fifth signal unless it can absorb the others — an ambitious unification that no single working group has yet achieved.
Whether RSL becomes a standard, a footnote, or a negotiating position borrowed by lawyers drafting bilateral data-licensing agreements, it names something real: the crawl has economic value, and the web has no agreed language for pricing it.



