Most scraping projects discover what protects a site the hard way: the prototype works on Monday, a challenge page appears on Tuesday, and by Friday the team is rebuilding around a headless browser. Much of that could be known on day one. Bot-management products leave visible traces in a site’s response headers and cookies, and a single ordinary request is enough to read them.
We built a short lookup that does exactly that, ran it against the homepages of the top 1,000 domains, and recorded what a plain HTTP client was served. This guide shares the results, the code, and how to use the answer when planning a project.
Key takeaways
- One request reveals a lot. Vendor cookies and headers identified a bot-protection or edge-security vendor on 37% of the 653 top-site homepages we could load, and an actively running bot-management product on 16%.
- Cloudflare was by far the most common, on a quarter of homepages, followed by Akamai. DataDome, AWS WAF, Imperva and HUMAN appeared far less often, and DataDome and HUMAN challenged or blocked every request we sent them.
- A plain HTTP client claiming to be Chrome was served 81% of homepages, challenged on 10% and blocked on 9%. With the library’s default User-Agent, blocks rose to 16%.
- Pretending to be a browser can backfire. One large retailer’s ten country storefronts served full homepages to the honest library User-Agent and a one- to two-kilobyte challenge page to the request that claimed to be Chrome.
- Treat the lookup as a planning input: it tells you what to expect, whether an official API or permission is the better route, and how to budget the project.
What a single request can reveal
Bot-management and edge-security products work by sitting in front of a website and inspecting every request. To do that, most of them set cookies on the visitor’s browser or add response headers, and those names are stable and often documented:
| Vendor | Typical evidence | Where it comes from |
|---|---|---|
| Cloudflare | cf-ray header on everything it serves; __cf_bm cookie where Bot Management or Bot Fight Mode is on; cf-mitigated: challenge on a challenge page | Cloudflare’s documentation |
| Akamai | Akamai-GRN request-number header; _abck, ak_bmsc, bm_sz cookies from Bot Manager | Akamai’s documentation for the header; site cookie policies for the cookies |
| DataDome | datadome cookie; x-datadome header | DataDome’s documentation |
| HUMAN | _px3, _pxhd, _pxvid cookies | HUMAN’s documentation |
| Imperva | visid_incap_, incap_ses_, nlbi_ cookies | Site cookie policies |
| AWS WAF | x-amzn-waf-action: challenge with HTTP 202 on a challenge; aws-waf-token cookie | AWS documentation |
Two distinctions matter when reading the results. Being behind a content delivery network is not the same as running bot management: a cf-ray header means the site uses Cloudflare, while a __cf_bm cookie means a Cloudflare bot product is actively scoring visitors. And a missing signature proves nothing; many sites run their own detection or a product that leaves no trace on the first response.
The code
The module below reads one response and returns the vendors it detects with the evidence, whether a bot-management product is active, and what the request got: served, challenged or blocked.
import re
# Cookies that each vendor's bot or WAF product sets, and headers it adds. Cookies are the strongest evidence.
SIGNATURES = {
"Cloudflare": {"headers": ["cf-ray"], "cookies": [r"^__cf_bm$", r"^cf_clearance$"]},
"Akamai": {"headers": ["akamai-grn"], "cookies": [r"^_abck$", r"^ak_bmsc$", r"^bm_sz$", r"^bm_sv$"]},
"DataDome": {"headers": ["x-datadome"], "cookies": [r"^datadome$"]},
"HUMAN": {"headers": [], "cookies": [r"^_px(3|hd|vid|cvid|de)$"]},
"Imperva": {"headers": [], "cookies": [r"^visid_incap_\d+$", r"^incap_ses_", r"^nlbi_\d+$", r"^reese84$"]},
"AWS WAF": {"headers": ["x-amzn-waf-action"], "cookies": [r"^aws-waf-token$"]},
}
# Evidence of an active bot-management product, as opposed to only the CDN in front of the site.
BOT_PRODUCT_COOKIES = [r"^__cf_bm$", r"^cf_clearance$", r"^_abck$", r"^ak_bmsc$", r"^bm_sz$", r"^datadome$", r"^_px", r"^reese84$"]
def cookie_names(response):
"""Names of every cookie the response tried to set."""
return {c.split("=", 1)[0].strip() for c in response.raw.headers.getlist("Set-Cookie")}
def detect_stack(response):
"""Return {vendor: [evidence]} from one response's headers and cookies."""
headers = {k.lower() for k in response.headers}
cookies = cookie_names(response)
found = {}
for vendor, sig in SIGNATURES.items():
evidence = [f"header {h}" for h in sig["headers"] if h in headers]
evidence += [f"cookie {c}" for c in sorted(cookies) if any(re.match(p, c) for p in sig["cookies"])]
if evidence:
found[vendor] = evidence
server = response.headers.get("server", "").lower()
if server == "cloudflare":
found.setdefault("Cloudflare", []).append("server header")
if "akamaighost" in server:
found.setdefault("Akamai", []).append("server header")
return found
def bot_product_active(response):
"""True when a cookie shows a bot-management product, not just a CDN, handled the request."""
return any(re.match(p, c) for c in cookie_names(response) for p in BOT_PRODUCT_COOKIES)
def outcome(response):
"""Classify what a request got: served, challenged or blocked."""
body = response.text[:20000].lower()
if (response.headers.get("cf-mitigated", "").lower() == "challenge"
or response.headers.get("x-amzn-waf-action", "").lower() == "challenge"
or "captcha-delivery.com" in body):
return "challenged"
if response.status_code in (401, 403, 405, 429) or response.status_code >= 500 or "incapsula incident id" in body:
return "blocked"
return "served"
It works on a response from the requests library and reads cookies from the raw Set-Cookie headers, so it sees cookies the client would not otherwise keep. Note what “served” means here: no challenge or block signal was found. It does not prove the page contained the real content, a gap we come back to below.
What we found on the top 1,000 domains
On 3 October 2026 we requested the homepage of each of the top 1,000 domains in the Tranco ranking, once with a Chrome User-Agent string and once with the requests library’s default, from a single connection in Romania. Both requests came from the same Python HTTP client, which runs no JavaScript and does not have a browser’s network fingerprint. 653 distinct sites returned an HTML homepage; the rest were content networks, API hosts and other domains without one.
| Vendor detected | Homepages | With an active bot-management cookie |
|---|---|---|
| Cloudflare | 162 (24.8%) | 79 |
| Akamai | 55 (8.4%) | 16 |
| AWS WAF | 14 (2.1%) | 0 |
| DataDome | 8 (1.2%) | 8 |
| HUMAN | 2 | 2 |
| Imperva | 2 | 0 |
| Any of the above | 241 (36.9%) | 105 (16.1%) |
| None detected | 412 (63.1%) | 0 |
Only two homepages showed two vendors at once. Ten of the 14 AWS WAF sites were country storefronts of the same retailer.
What the requests got:
| Request | Served | Challenged | Blocked |
|---|---|---|---|
| Chrome User-Agent, 653 homepages | 530 (81.2%) | 64 (9.8%) | 59 (9.0%) |
| Library default User-Agent, 650 homepages | 500 (76.9%) | 49 (7.5%) | 101 (15.5%) |
And broken down by what protected the site, counting sites where both requests completed:
| Protection detected | Sites | Chrome User-Agent: challenged or blocked | Library default: challenged or blocked |
|---|---|---|---|
| Cloudflare | 162 | 54 | 63 |
| Akamai | 54 | 26 | 22 |
| DataDome | 7 | 7 | 7 |
| None detected | 411 | 19 | 53 |
What the results say
Most of the top web answers a plain request. Four in five homepages served a basic HTTP client claiming to be Chrome without a visible challenge. A homepage is the easiest page on a site, though; search results, prices and checkout flows are usually protected more tightly than the front door.
Simple filters and serious products behave differently. On sites with no detectable vendor, the library’s default User-Agent was challenged or blocked almost three times as often as the Chrome string, 53 times against 19. That is the signature of simple User-Agent filtering. On sites running DataDome, both requests were challenged or blocked every time; the User-Agent made no difference.
A mismatched identity can be worse than an honest one. The retailer’s ten country storefronts answered the request claiming to be Chrome with a one- to two-kilobyte challenge or placeholder page, and the honest library User-Agent with the full homepage, between 678 KB and 1.4 MB. We repeated the check for all ten and the pattern held. A Chrome User-Agent string arriving on a connection that does not look like Chrome, as explained in TLS and HTTP/2 fingerprinting, is itself a signal.
A successful status is not a successful page. Three of those placeholder pages came back with HTTP 200 and no challenge header, so a status-only check counts them as served. Validate content, not just status codes, as covered in the silent failure rate and how to tell when a site is serving you fake or blocked content.
Using the lookup to plan a project
Run the lookup on the pages you actually need, not just the homepage, and let the answer shape the plan:
| What the lookup shows | What it usually means | Sensible next step |
|---|---|---|
| No vendor, plain request served | Little or no bot management on that page | Plain HTTP with honest identification, modest pace, and the right headers |
| CDN only, no bot cookie | Edge security without active bot scoring | Plain HTTP, but watch for challenges as volume grows |
| Bot-management cookie, request served | Active scoring that allowed this request | Expect challenges at scale; monitor with a target health score |
| Challenged on the first request | A deliberate decision to filter automated clients | Look for an official API or data feed first, then consider whether a browser is justified, per when you need a headless browser |
| Blocked on the first request | Strict rules, often by network or region | Check whether the data is available another way, including the API behind the page, and whether the block is regional |
The lookup also says something about intent. A site that runs an active bot-management product has decided to control automated access, and that is worth weighing alongside its terms and its robots.txt, as discussed in robots.txt, AI opt-outs and reservation signals. For many projects, the right response to a heavily protected target is a licence, a partnership or an official API rather than an escalation. Where collection is appropriate, the escalation ladder for protected sites explains the options in order of cost.
Limits of the measurement
- Homepages only. Deeper pages are often protected differently.
- One request each, from one network. Results can differ by country, network reputation, time of day and request history; we covered regional differences in which countries get geo-blocked most.
- No browser. Many challenges are designed to be passed by a real browser running JavaScript; our client could not, by design.
- Signatures miss things. In-house detection and products that leave no first-request trace are counted as “none detected”, so the true share of protected sites is higher.
The bottom line
A single ordinary request tells you most of what you need to plan a scraping project: who protects the site, whether bot scoring is active, and how a plain client is treated. On the top 1,000 domains, most homepages answered, a quarter sat behind Cloudflare, one in six ran an active bot-management product, and the strictest products challenged or blocked every request we sent.
Run the lookup before writing the scraper, on the pages you need. Let it tell you when to keep things simple, when to plan for a browser, and when to look for an official route to the data instead.
Sources and references
- Cloudflare, Cloudflare cookies and detecting a challenge page response.
- Akamai, Global request number.
- DataDome, Cookies and stored data.
- HUMAN, Use of cookies and web storage.
- AWS, CAPTCHA and Challenge in AWS WAF.
- Tranco, list Q2K34.
- Homepages requested by Shifter on 3 October 2026, using the code above.