robots.txt as convention, GPTBot and Google-Extended, the Data Provenance Initiative's measurement, Cloudflare's 2025 default, and the licensing deals signed instead.
The legal and technical status of robots.txt — a 1994 convention with no enforcement mechanism, whose moral weight depends entirely on the compliance of the party it addresses.
OpenAI's August 2023 GPTBot disclosure, Google-Extended's arrival weeks later, and the wave of robots.txt blocks that followed — measured by the Data Provenance Initiative and documented in Cloudflare's crawl data.
What the Data Provenance Initiative is, how it measured the rate of web blocking against AI crawlers in 2023–2024, and what the numbers showed about which categories of publisher blocked first and fastest.
Cloudflare's decision to make AI bot blocking a default setting for new customers — the announcement, what it blocks, and what it means for the fraction of the web that runs behind Cloudflare's network.
The Responsible Scraping License as a proposed standard for machine-readable content licensing — what it specifies, who is pushing it, and whether a voluntary licensing framework can resolve a conflict that robots.txt was never designed to handle.