For most of its thirty-year life, robots.txt was a courtesy: a plain text file asking crawlers to stay out of certain paths, with nothing behind it but good manners. That has changed. It became a formal internet standard in 2022. In the EU, machine-readable opt-outs now decide whether text and data mining is lawful, and AI model providers have a legal duty to find and respect them. Meanwhile a crop of new signals has appeared, each claiming to say what a site allows AI systems to do with its content.
For anyone collecting public web data, and especially for anyone whose data ends up training or feeding AI models, the practical question is simple: which of these signals do you have to read, what do they mean, and what should you do about them? This guide sets out the current state, as of September 2026.
Key takeaways
- robots.txt is an IETF standard, RFC 9309. It defines precise matching rules, but it states plainly that its rules “are not a form of access authorization.”
- In the EU, the commercial text and data mining exception applies only where rightholders have not reserved their rights “in an appropriate manner, such as machine-readable means” for content online. German courts have held that such a reservation is only effective in machine-readable form.
- General-purpose AI model providers have been required since 2 August 2025 to identify and comply with those reservations. The Commission’s power to fine them applies from 2 August 2026.
- New preference signals exist, from Cloudflare’s Content Signals to the IETF’s draft AI preferences vocabulary, but none is yet a finished standard. Record them even where they are not binding.
- Using a proxy network does not change what a site has asked for. The obligations attach to what you collect and how you use it, not to the IP address you collect it from.
robots.txt is a standard now
RFC 9309, published in September 2022, turned a de facto convention into a Proposed Standard. Several of its rules surprise people who have only ever read robots.txt files by eye:
- The most specific rule wins. “The most specific match found MUST be used,” measured by the length of the matching rule, with
allowpreferred when an allow and a disallow rule are equivalent. Order in the file does not matter. - Agent matching is case-insensitive, and a crawler with no matching group falls back to the
*group. - Errors mean different things. If robots.txt is unavailable, with a 4xx response, a crawler may access anything. If it is unreachable because of server or network errors, the crawler “MUST assume complete disallow.”
- Caching is bounded. Crawlers “SHOULD NOT use the cached version for more than 24 hours, unless the robots.txt file is unreachable.”
And the limit that matters most: “These rules are not a form of access authorization.” robots.txt does not make content private, and ignoring it is not the same as breaking into a system. What gives it weight in some contexts is law, not the protocol itself.
Why it matters legally in the EU
The EU’s copyright rules allow text and data mining of lawfully accessible content, with two different exceptions. One covers research organisations and cultural heritage institutions doing scientific research, and has no opt-out. The other covers everyone else, including commercial mining, but applies only where the use “has not been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online.”
The AI Act then applies that rule to AI. Providers of general-purpose AI models must “put in place a policy to comply with Union law on copyright and related rights, and in particular to identify and comply with, including through state-of-the-art technologies, a reservation of rights” made under the copyright rules. Those obligations have applied since 2 August 2025. The Commission’s fining powers over general-purpose AI providers, up to 3% of worldwide turnover, apply from 2 August 2026, and models placed on the market before August 2025 have until 2 August 2027 to comply.
The voluntary General-Purpose AI Code of Practice, published in July 2025, makes this concrete. Signatories commit “to employ web-crawlers that read and follow instructions expressed in accordance with the Robot Exclusion Protocol (robots.txt), as specified in” RFC 9309, and to identify and comply with other appropriate machine-readable protocols. Two caveats in the Code are worth noting. Adherence “does not constitute compliance with Union law on copyright and related rights.” And the commitment does not affect how copyright applies to content “scraped or crawled from the internet by third parties” and used by signatories. Buying a dataset does not launder the reservations that applied to its sources.
What courts have said
Case law is still forming, and two decisions illustrate the direction.
In Germany, a photographer sued the non-profit LAION over the use of his image in an AI training dataset. The Hamburg courts dismissed the claim. On appeal, the Higher Regional Court held, in the words of the Federal Court of Justice’s summary, that a reservation for online works is only effective “in maschinenlesbarer Form”, in machine-readable form, and that the claimant had not shown that his reservation, written in natural language in terms of use, was machine-readable at the relevant time. The Federal Court of Justice heard the further appeal on 3 September 2026 and has scheduled its decision for 17 December 2026.
In the Netherlands, an Amsterdam court ruled on 30 October 2024 in a dispute between newspaper publishers and a news aggregator. The aggregator argued, and the publishers did not dispute, that their robots.txt excluded only certain AI bots, such as GPTBot, ChatGPT-User and CCBot. That was not enough to establish that rights had been reserved in an appropriate machine-readable way for the aggregator’s use, so the exception applied.
Both point the same way: machine-readable signals are what count, and a signal aimed at specific bots may not reserve rights against everyone else. Neither is final guidance, and both depend on facts. Take specific questions to counsel.
The signals you will meet
| Signal | Where it lives | Status | What it expresses |
|---|---|---|---|
| robots.txt Allow and Disallow | /robots.txt | IETF standard, RFC 9309 | Whether a named crawler may access a path |
| Content Signals | Lines in robots.txt, such as Content-Signal: search=yes, ai-train=no | Cloudflare policy, published September 2025 | Preferences for search, AI input and AI training |
| AI preferences (Content-Usage) | robots.txt lines or an HTTP header, such as Content-Usage: train-ai=n | IETF working group drafts, not yet standards | Preferences for AI training, AI use and search |
| TDMRep | /.well-known/tdmrep.json, a tdm-reservation HTTP header or HTML meta tag | W3C Community Group report, not a W3C Standard | Whether text and data mining rights are reserved |
| Terms of service | The site’s legal pages | Contract and, in the EU, possibly a reservation if machine-readable | Whatever the site’s terms say |
Two details matter. Cloudflare’s Content Signals define search, ai-input (for retrieval, grounding and generated answers) and ai-train, and the policy states that “content signals express preferences; they are not technical countermeasures against scraping.” It also declares that restrictions expressed through the signals are express reservations of rights under the EU copyright rules. The IETF drafts, meanwhile, are genuinely unfinished: the current vocabulary draft carries a note that its contents “DO NOT REFLECT CONSENSUS of the Working Group”. Expect the labels to change.
How the major AI crawlers treat them
The operators publish their own rules, and they differ, especially for fetches made on behalf of a user.
| Operator | Agent | What blocking it does, per the operator |
|---|---|---|
| OpenAI | GPTBot | Indicates content “should not be used in training generative AI foundation models” |
| OpenAI | OAI-SearchBot | The site is not shown in ChatGPT search answers, though it can still appear as navigational links |
| OpenAI | ChatGPT-User | User-initiated; “robots.txt rules may not apply” |
| Google-Extended | A robots.txt token with no separate user agent; controls use for training and grounding Gemini models, and “does not impact a site’s inclusion in Google Search” | |
| User-triggered fetchers | ”generally ignore robots.txt rules” | |
| Anthropic | ClaudeBot, Claude-User, Claude-SearchBot | Blocking ClaudeBot excludes future content from training; blocking Claude-User prevents retrieval for user queries; the bots honour robots.txt and crawl-delay |
For a data team, the lesson is that “blocking AI” is not one switch. A site can allow search indexing while refusing training, or refuse training while still being fetched live when a user asks about it. Read the signals for the use you actually have.
How many sites opt out
Opt-outs have grown quickly. The Data Provenance Initiative’s “Consent in Crisis” study, an audit of 14,000 web domains published in July 2024, found that in a single year robots.txt restrictions made “~5%+ of all tokens in C4, or 28%+ of the most actively maintained, critical sources in C4, fully restricted,” and that “for Terms of Service crawling restrictions, a full 45% of C4 is now restricted.” Coverage is uneven at the top of the web, though: Cloudflare reported in July 2025 that “only about 37% of the top 10,000 domains currently have a robots.txt file,” and that GPTBot was disallowed in 7.8% of those files.
What a responsible collector should do
Whether or not you train models, a clear policy protects you and the sites you rely on.
- Fetch and obey robots.txt properly. Cache it for no more than 24 hours, treat server errors as “disallow everything”, and match rules the RFC 9309 way.
- Know your purpose. Indexing, retrieval for live answers and model training are different uses, and the newer signals treat them differently. Decide which your collection serves, and read the signals for that use.
- Record the signals with the data. Store the robots.txt rules, content signals and TDMRep headers that applied at collection time alongside each record. If a dataset is later used for AI, that record is how you show reservations were identified. It belongs in the same provenance record as the vantage point and capture time.
- Assess terms of service with counsel. They can matter contractually everywhere, and in the EU they may count as a reservation if expressed in a machine-readable way.
- Pace your collection. Honouring crawl limits and backing off under load is part of the same good faith, as covered in rate limiting and request throttling.
A caution on tooling: Python’s built-in urllib.robotparser applies rules in file order rather than by most specific match. Given Disallow: /shop followed by Allow: /shop/public, it reports /shop/public/item as disallowed; swap the two lines and it says allowed. RFC 9309 says allowed either way. A small checker that follows the RFC, and records the AI preference lines it finds, looks like this:
import re
import urllib.error
import urllib.request
def _groups(text):
"""Parse robots.txt into (agents, rules, extras) groups, following RFC 9309 grouping."""
groups, agents, rules, extras, last = [], [], [], [], None
for raw in text.splitlines():
line = raw.split("#", 1)[0].strip()
if ":" not in line:
continue
key, value = (p.strip() for p in line.split(":", 1))
key = key.lower()
if key == "user-agent":
if last != "user-agent" and agents:
groups.append((agents, rules, extras))
agents, rules, extras = [], [], []
agents.append(value.lower())
elif key in ("allow", "disallow") and agents:
rules.append((key, value))
elif key in ("content-signal", "content-usage") and agents:
extras.append((key, value))
last = key
if agents:
groups.append((agents, rules, extras))
return groups
def _pattern(path):
regex = re.escape(path).replace(r"\*", ".*")
return re.compile(regex[:-2] + "$" if regex.endswith(r"\$") else regex)
def check(robots_txt, user_agent, path):
"""Allow/disallow per RFC 9309 (most specific match wins, allow on ties), plus AI signals."""
token = user_agent.lower()
groups = _groups(robots_txt)
matched = [g for g in groups if token in g[0]] or [g for g in groups if "*" in g[0]]
rules = [r for g in matched for r in g[1]]
extras = [e for g in matched for e in g[2]]
best = None
for kind, value in rules:
if value and _pattern(value).match(path):
length = len(value)
if best is None or length > best[1] or (length == best[1] and kind == "allow"):
best = (kind, length)
return {"allowed": best is None or best[0] == "allow", "signals": extras}
def fetch_robots(origin, opener=None):
"""Fetch robots.txt. Per RFC 9309: 4xx means no rules; 5xx or network failure means disallow all."""
opener = opener or urllib.request.build_opener()
try:
with opener.open(origin.rstrip("/") + "/robots.txt", timeout=30) as r:
return r.read(512 * 1024).decode("utf-8", "replace")
except urllib.error.HTTPError as e:
return "" if 400 <= e.code < 500 else "User-agent: *\nDisallow: /"
except OSError:
return "User-agent: *\nDisallow: /"
It is deliberately small. Production crawlers should also check TDMRep headers and the /.well-known/tdmrep.json file where training is in scope, and keep a dated copy of every robots.txt they relied on.
Where proxies fit
Residential proxies change where a request appears to come from. They do not change what a site has asked of you, and they do not change your obligations under copyright law or a site’s terms. The EU rules attach to the use of the content, and the Code of Practice explicitly keeps third-party collection in scope. Use a proxy network to see the web as real users in a given market see it, to spread load politely and to measure honestly, and apply the same robots.txt and reservation checks you would apply from your own servers. The broader case is made in ethical residential proxies for AI data collection.
The bottom line
robots.txt has grown up. It is a standard with precise rules, and in the EU, machine-readable reservations built on it and alongside it now decide whether text and data mining is permitted, with AI providers under a direct legal duty to honour them and fines available from August 2026. New signals for search, AI input and training are spreading fast, even though the standards behind them are unfinished.
For a data team, the workable policy is the same regardless of how the details settle: parse robots.txt correctly, know which use your collection serves, record every signal that applied at the moment of collection, and treat preferences you are not strictly bound by as information worth keeping rather than noise to discard. The teams that do this now will not have to reconstruct it later.
Sources and references
- IETF, RFC 9309: Robots Exclusion Protocol, September 2022.
- Directive (EU) 2019/790 on copyright and related rights in the Digital Single Market. Text and data mining exceptions.
- Regulation (EU) 2024/1689, the AI Act. Obligations for providers of general-purpose AI models and dates of application.
- European Commission, General-Purpose AI Code of Practice, published 10 July 2025, copyright chapter.
- Bundesgerichtshof, press release 085/2026 on I ZR 281/25, and hearing and decision dates.
- Rechtbank Amsterdam, judgment of 30 October 2024, ECLI:NL:RBAMS:2024:6563.
- IETF AI Preferences working group, draft-ietf-aipref-vocab and draft-ietf-aipref-attach.
- W3C Community Group, TDM Reservation Protocol, 10 May 2024.
- Cloudflare blog, Content Signals Policy, 24 September 2025, and Reid Tatoris, on controlling content use for AI training, 1 July 2025.
- OpenAI, Overview of OpenAI crawlers; Google, common crawlers and user-triggered fetchers; Anthropic, crawler help article.
- Longpre et al., Consent in Crisis: The Rapid Decline of the AI Data Commons, July 2024.
This article is general information, not legal advice. Consult counsel about your obligations in each jurisdiction.