AI companies send their own crawlers across the web, and site owners answer through robots.txt, the plain-text file that tells automated visitors what they may fetch. Each major AI company publishes the names its crawlers answer to, so a site can block one, some or all of them while still welcoming search engines.
We wanted a current picture of who does. On 6 October 2026 we requested the robots.txt of every one of the top 10,000 domains in the Tranco ranking, evaluated the rules for 16 AI crawlers and two search crawlers, and recorded which of them each site lets onto its homepage. This guide shares the results, the evaluator we wrote, and what they mean for anyone collecting web data.
Key takeaways
- 5,754 of the top 10,000 domains served a robots.txt we could read. One in five of those, 20.4%, blocks at least one AI crawler from its homepage.
- The most blocked crawlers are Common Crawl’s CCBot (16.2%), ByteDance’s Bytespider (15.3%), OpenAI’s GPTBot (15.1%) and Anthropic’s ClaudeBot (14.0%).
- Search engines are almost never blocked: Googlebot by 2.6% and Bingbot by 3.3%, most of them sites that block every crawler. 17.3% of sites let both search engines in while blocking at least one AI crawler.
- Many sites separate training from answering. Of the 871 sites blocking GPTBot, 345 still allow ChatGPT-User, which fetches pages when a person asks ChatGPT about them.
- The bigger the site, the more likely it blocks: 29.5% of the top 1,000 block at least one AI crawler, against 18.8% of those ranked 5,001 to 10,000.
How we measured
For each domain we requested https://<domain>/robots.txt, falling back to the www. host, with a User-Agent that named us and linked to our site, at a modest pace. We then evaluated the rules for each crawler against the homepage path /.
We followed RFC 9309, the robots.txt standard, rather than the shortcut most libraries take:
- A crawler uses every group that names it, and falls back to the
*group only when no group does. - The longest matching rule wins, and
Allowwins a tie. - A 4xx response means there are no rules; a 5xx response means everything is disallowed.
Python’s built-in parser applies the first matching rule instead of the longest, which misreads common files such as Disallow: / followed by Allow: /$, so we wrote our own evaluator and tested it against the standard’s cases. We counted a crawler as “blocked” when it may not fetch the homepage, and recorded separately whether the file names the crawler or blocks it only through *.
Of the 10,000 domains, 5,877 returned a parseable robots.txt and 5,754 of those were distinct sites; the rest were mostly API hosts, content delivery networks and other domains that serve no website. 460 answered with an HTML page instead of a robots file, which the standard treats as no rules.
What we found
| Crawler | Operator | Purpose | Blocked from homepage | Named in the file |
|---|---|---|---|---|
| CCBot | Common Crawl | Open web archive widely used for AI training | 16.2% | 15.2% |
| Bytespider | ByteDance | Crawling for AI models | 15.3% | 13.1% |
| GPTBot | OpenAI | Training | 15.1% | 17.7% |
| ClaudeBot | Anthropic | Training | 14.0% | 15.1% |
| meta-externalagent | Meta | Training and AI products | 12.5% | 10.8% |
| Google-Extended | Training and grounding Gemini, separate from Search | 12.3% | 13.5% | |
| anthropic-ai | Anthropic | Older Anthropic token | 11.7% | 9.9% |
| Applebot-Extended | Apple | Training, separate from Applebot | 11.7% | 10.1% |
| Amazonbot | Amazon | Crawling for Amazon services | 11.7% | 10.3% |
| cohere-ai | Cohere | AI products | 11.7% | 9.2% |
| PerplexityBot | Perplexity | Search index for answers | 10.8% | 12.7% |
| ChatGPT-User | OpenAI | Fetches a page when a user asks | 9.4% | 11.7% |
| Perplexity-User | Perplexity | Fetches a page when a user asks | 7.7% | 6.2% |
| OAI-SearchBot | OpenAI | Search results in ChatGPT | 7.5% | 9.9% |
| Claude-User | Anthropic | Fetches a page when a user asks | 7.3% | 6.1% |
| Claude-SearchBot | Anthropic | Search results in Claude | 7.2% | 6.1% |
| Bingbot | Microsoft | Web search | 3.3% | 7.4% |
| Googlebot | Web search | 2.6% | 9.1% |
Percentages are of the 5,754 sites. “Named” counts files that address the crawler by name, whether to block it or to allow it, which is why the two columns differ: some sites name a crawler only to welcome it, and 153 sites block every crawler through * without naming any.
Three patterns
AI crawlers are blocked; search engines are not. 17.3% of sites let both Googlebot and Bingbot onto the homepage while blocking at least one AI crawler. Only 0.2% block Googlebot by name. Site owners are drawing a clear line between being indexed for search and being collected for AI.
Training is blocked more than answering. The crawlers that gather training data are blocked noticeably more often than the ones that fetch a page because a person asked about it: GPTBot on 15.1% of sites against 9.4% for ChatGPT-User, ClaudeBot on 14.0% against 7.3% for Claude-User. Of the 871 sites that block GPTBot, 345 still allow ChatGPT-User, and of the 806 that block ClaudeBot, 390 still allow Claude-User. Many sites, in other words, want to appear in AI answers without contributing to training.
Blocking rises with size. Among the top 1,000 sites, 29.5% block at least one AI crawler and 21.0% block GPTBot. Among sites ranked 5,001 to 10,000, those figures fall to 18.8% and 13.9%. Large publishers and platforms, whose content is most valuable for training, are the ones most likely to opt out.
Only 5.2% of sites block all five of the most prominent AI crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot and PerplexityBot) while still allowing Googlebot. Most opt-outs are selective: a site blocks some companies and not others, or blocks training and allows answering.
The code
The module below parses a robots.txt file and evaluates it for any crawler under RFC 9309’s rules. It has no dependencies.
import re
from urllib.parse import quote, unquote
def parse_groups(text):
"""Split robots.txt into groups: a list of (user-agent tokens, [(allow, path)])."""
groups, agents, rules, in_rules = [], [], [], False
for raw in text.splitlines():
line = raw.split("#", 1)[0].strip()
if ":" not in line:
continue
key, value = (part.strip() for part in line.split(":", 1))
key = key.lower()
if key == "user-agent":
if in_rules: # a user-agent line after rules starts a new group
groups.append((agents, rules))
agents, rules, in_rules = [], [], False
agents.append(value.lower())
elif key in ("allow", "disallow") and agents:
in_rules = True
if value: # an empty Disallow allows everything
rules.append((key == "allow", value))
if agents:
groups.append((agents, rules))
return groups
def rules_for(groups, product_token):
"""RFC 9309: use every group naming the crawler's product token; otherwise the '*' groups."""
token = product_token.lower()
named = [r for agents, rules in groups for r in rules if token in agents]
if any(token in agents for agents, _ in groups):
return named
return [r for agents, rules in groups for r in rules if "*" in agents]
def _pattern(path):
regex = "".join(".*" if ch == "*" else re.escape(ch) for ch in path.rstrip("$"))
return re.compile(regex + ("$" if path.endswith("$") else ""))
def allowed(groups, product_token, path="/"):
"""Longest matching rule wins; Allow wins a tie (RFC 9309, section 2.2.2)."""
target = quote(unquote(path), safe="/?=&*$%:@!,;~+")
best_len, verdict = -1, True
for allow, rule in rules_for(groups, product_token):
rule = quote(unquote(rule), safe="/?=&*$%:@!,;~+")
if _pattern(rule).match(target):
length = len(rule)
if length > best_len or (length == best_len and allow):
best_len, verdict = length, allow
return verdict
def names(groups, product_token):
"""Does the file address this crawler by name, rather than only through '*'?"""
return any(product_token.lower() in agents for agents, _ in groups)
And checking a site, here our own:
import requests
from robots import parse_groups, allowed, names
text = requests.get("https://shifter.io/robots.txt", timeout=15,
headers={"User-Agent": "ExampleSurvey/1.0 (+https://example.com/bot)"}).text
groups = parse_groups(text)
for crawler in ["GPTBot", "ClaudeBot", "Google-Extended", "CCBot", "Googlebot"]:
print(f"{crawler:16} homepage allowed: {allowed(groups, crawler, '/')} named: {names(groups, crawler)}")
Our file names the major AI crawlers and allows them explicitly, so all five come back allowed, and Googlebot is allowed through *. We tested the evaluator against 16 cases, including longest-match precedence, the tie rule, wildcard and end-of-path patterns, merged groups and empty files.
What it means for data collection
- Read robots.txt per crawler, not just for
*. A site that welcomes everyone through*may still block a named crawler, and a growing share of the web now does exactly that. - Identify your collector honestly. Robots rules only work if crawlers say who they are. If you collect under your own name, check the rules for that name and for the purpose you are collecting for.
- Respect the distinction sites are drawing. Many sites now separate search, training and on-demand fetching. A collector gathering training data should treat a block on training crawlers as applying to it, even if its own name is not listed.
- Recheck regularly. Opt-outs change as sites update their files; a list built six months ago will be out of date.
Our guide to robots.txt, AI opt-outs and reservation signals covers the other signals sites use and the legal context in Europe, and web scraping best practices covers collecting without harming the sites you visit. Training data pipelines that need to honour opt-outs at fetch time are covered in building large-scale training datasets.
Limits of the measurement
- Homepage only. We evaluated the path
/. Many sites block AI crawlers from specific sections, such as articles or search pages, while leaving the homepage open, so the share of sites restricting AI crawlers somewhere is higher than these figures. - robots.txt only. Sites also signal through HTTP headers, meta tags and terms of service, and enforce through bot management. A site that does not block a crawler in robots.txt may still block it at the network level.
- One snapshot. These are the files as served on 6 October 2026, from one network.
- Top sites only. The top 10,000 domains are not the whole web; smaller sites may behave differently.
The bottom line
One top site in five now tells at least one AI crawler to stay away from its homepage, while almost all of them still welcome search engines. Opt-outs are selective: training crawlers are blocked more often than the ones that fetch pages for users, and the largest sites block the most.
For anyone collecting web data, the practical lesson is to read robots.txt for the specific crawler and purpose involved, with an evaluator that follows the standard, and to recheck it as the web keeps redrawing these lines.
Sources and references
- IETF, RFC 9309: Robots Exclusion Protocol, September 2022.
- Tranco, list Q2K34.
- robots.txt files of the top 10,000 domains, requested by Shifter on 6 October 2026 using the code above.