Scraping

Is Your Target Blockable? An Anti-Bot Stack Lookup

Before building a scraper, find out what protects the target. We checked the top 1,000 domains: which anti-bot vendors they use and how a plain client fared.

Chris Collins

Chris Collins

October 3, 2026 · 10 min read

Most scraping projects discover what protects a site the hard way: the prototype works on Monday, a challenge page appears on Tuesday, and by Friday the team is rebuilding around a headless browser. Much of that could be known on day one. Bot-management products leave visible traces in a site’s response headers and cookies, and a single ordinary request is enough to read them.

We built a short lookup that does exactly that, ran it against the homepages of the top 1,000 domains, and recorded what a plain HTTP client was served. This guide shares the results, the code, and how to use the answer when planning a project.

Key takeaways

  • One request reveals a lot. Vendor cookies and headers identified a bot-protection or edge-security vendor on 37% of the 653 top-site homepages we could load, and an actively running bot-management product on 16%.
  • Cloudflare was by far the most common, on a quarter of homepages, followed by Akamai. DataDome, AWS WAF, Imperva and HUMAN appeared far less often, and DataDome and HUMAN challenged or blocked every request we sent them.
  • A plain HTTP client claiming to be Chrome was served 81% of homepages, challenged on 10% and blocked on 9%. With the library’s default User-Agent, blocks rose to 16%.
  • Pretending to be a browser can backfire. One large retailer’s ten country storefronts served full homepages to the honest library User-Agent and a one- to two-kilobyte challenge page to the request that claimed to be Chrome.
  • Treat the lookup as a planning input: it tells you what to expect, whether an official API or permission is the better route, and how to budget the project.

What a single request can reveal

Bot-management and edge-security products work by sitting in front of a website and inspecting every request. To do that, most of them set cookies on the visitor’s browser or add response headers, and those names are stable and often documented:

VendorTypical evidenceWhere it comes from
Cloudflarecf-ray header on everything it serves; __cf_bm cookie where Bot Management or Bot Fight Mode is on; cf-mitigated: challenge on a challenge pageCloudflare’s documentation
AkamaiAkamai-GRN request-number header; _abck, ak_bmsc, bm_sz cookies from Bot ManagerAkamai’s documentation for the header; site cookie policies for the cookies
DataDomedatadome cookie; x-datadome headerDataDome’s documentation
HUMAN_px3, _pxhd, _pxvid cookiesHUMAN’s documentation
Impervavisid_incap_, incap_ses_, nlbi_ cookiesSite cookie policies
AWS WAFx-amzn-waf-action: challenge with HTTP 202 on a challenge; aws-waf-token cookieAWS documentation

Two distinctions matter when reading the results. Being behind a content delivery network is not the same as running bot management: a cf-ray header means the site uses Cloudflare, while a __cf_bm cookie means a Cloudflare bot product is actively scoring visitors. And a missing signature proves nothing; many sites run their own detection or a product that leaves no trace on the first response.

The code

The module below reads one response and returns the vendors it detects with the evidence, whether a bot-management product is active, and what the request got: served, challenged or blocked.

import re

# Cookies that each vendor's bot or WAF product sets, and headers it adds. Cookies are the strongest evidence.
SIGNATURES = {
    "Cloudflare": {"headers": ["cf-ray"], "cookies": [r"^__cf_bm$", r"^cf_clearance$"]},
    "Akamai":     {"headers": ["akamai-grn"], "cookies": [r"^_abck$", r"^ak_bmsc$", r"^bm_sz$", r"^bm_sv$"]},
    "DataDome":   {"headers": ["x-datadome"], "cookies": [r"^datadome$"]},
    "HUMAN":      {"headers": [], "cookies": [r"^_px(3|hd|vid|cvid|de)$"]},
    "Imperva":    {"headers": [], "cookies": [r"^visid_incap_\d+$", r"^incap_ses_", r"^nlbi_\d+$", r"^reese84$"]},
    "AWS WAF":    {"headers": ["x-amzn-waf-action"], "cookies": [r"^aws-waf-token$"]},
}
# Evidence of an active bot-management product, as opposed to only the CDN in front of the site.
BOT_PRODUCT_COOKIES = [r"^__cf_bm$", r"^cf_clearance$", r"^_abck$", r"^ak_bmsc$", r"^bm_sz$", r"^datadome$", r"^_px", r"^reese84$"]


def cookie_names(response):
    """Names of every cookie the response tried to set."""
    return {c.split("=", 1)[0].strip() for c in response.raw.headers.getlist("Set-Cookie")}


def detect_stack(response):
    """Return {vendor: [evidence]} from one response's headers and cookies."""
    headers = {k.lower() for k in response.headers}
    cookies = cookie_names(response)
    found = {}
    for vendor, sig in SIGNATURES.items():
        evidence = [f"header {h}" for h in sig["headers"] if h in headers]
        evidence += [f"cookie {c}" for c in sorted(cookies) if any(re.match(p, c) for p in sig["cookies"])]
        if evidence:
            found[vendor] = evidence
    server = response.headers.get("server", "").lower()
    if server == "cloudflare":
        found.setdefault("Cloudflare", []).append("server header")
    if "akamaighost" in server:
        found.setdefault("Akamai", []).append("server header")
    return found


def bot_product_active(response):
    """True when a cookie shows a bot-management product, not just a CDN, handled the request."""
    return any(re.match(p, c) for c in cookie_names(response) for p in BOT_PRODUCT_COOKIES)


def outcome(response):
    """Classify what a request got: served, challenged or blocked."""
    body = response.text[:20000].lower()
    if (response.headers.get("cf-mitigated", "").lower() == "challenge"
            or response.headers.get("x-amzn-waf-action", "").lower() == "challenge"
            or "captcha-delivery.com" in body):
        return "challenged"
    if response.status_code in (401, 403, 405, 429) or response.status_code >= 500 or "incapsula incident id" in body:
        return "blocked"
    return "served"

It works on a response from the requests library and reads cookies from the raw Set-Cookie headers, so it sees cookies the client would not otherwise keep. Note what “served” means here: no challenge or block signal was found. It does not prove the page contained the real content, a gap we come back to below.

What we found on the top 1,000 domains

On 3 October 2026 we requested the homepage of each of the top 1,000 domains in the Tranco ranking, once with a Chrome User-Agent string and once with the requests library’s default, from a single connection in Romania. Both requests came from the same Python HTTP client, which runs no JavaScript and does not have a browser’s network fingerprint. 653 distinct sites returned an HTML homepage; the rest were content networks, API hosts and other domains without one.

Vendor detectedHomepagesWith an active bot-management cookie
Cloudflare162 (24.8%)79
Akamai55 (8.4%)16
AWS WAF14 (2.1%)0
DataDome8 (1.2%)8
HUMAN22
Imperva20
Any of the above241 (36.9%)105 (16.1%)
None detected412 (63.1%)0

Only two homepages showed two vendors at once. Ten of the 14 AWS WAF sites were country storefronts of the same retailer.

What the requests got:

RequestServedChallengedBlocked
Chrome User-Agent, 653 homepages530 (81.2%)64 (9.8%)59 (9.0%)
Library default User-Agent, 650 homepages500 (76.9%)49 (7.5%)101 (15.5%)

And broken down by what protected the site, counting sites where both requests completed:

Protection detectedSitesChrome User-Agent: challenged or blockedLibrary default: challenged or blocked
Cloudflare1625463
Akamai542622
DataDome777
None detected4111953

What the results say

Most of the top web answers a plain request. Four in five homepages served a basic HTTP client claiming to be Chrome without a visible challenge. A homepage is the easiest page on a site, though; search results, prices and checkout flows are usually protected more tightly than the front door.

Simple filters and serious products behave differently. On sites with no detectable vendor, the library’s default User-Agent was challenged or blocked almost three times as often as the Chrome string, 53 times against 19. That is the signature of simple User-Agent filtering. On sites running DataDome, both requests were challenged or blocked every time; the User-Agent made no difference.

A mismatched identity can be worse than an honest one. The retailer’s ten country storefronts answered the request claiming to be Chrome with a one- to two-kilobyte challenge or placeholder page, and the honest library User-Agent with the full homepage, between 678 KB and 1.4 MB. We repeated the check for all ten and the pattern held. A Chrome User-Agent string arriving on a connection that does not look like Chrome, as explained in TLS and HTTP/2 fingerprinting, is itself a signal.

A successful status is not a successful page. Three of those placeholder pages came back with HTTP 200 and no challenge header, so a status-only check counts them as served. Validate content, not just status codes, as covered in the silent failure rate and how to tell when a site is serving you fake or blocked content.

Using the lookup to plan a project

Run the lookup on the pages you actually need, not just the homepage, and let the answer shape the plan:

What the lookup showsWhat it usually meansSensible next step
No vendor, plain request servedLittle or no bot management on that pagePlain HTTP with honest identification, modest pace, and the right headers
CDN only, no bot cookieEdge security without active bot scoringPlain HTTP, but watch for challenges as volume grows
Bot-management cookie, request servedActive scoring that allowed this requestExpect challenges at scale; monitor with a target health score
Challenged on the first requestA deliberate decision to filter automated clientsLook for an official API or data feed first, then consider whether a browser is justified, per when you need a headless browser
Blocked on the first requestStrict rules, often by network or regionCheck whether the data is available another way, including the API behind the page, and whether the block is regional

The lookup also says something about intent. A site that runs an active bot-management product has decided to control automated access, and that is worth weighing alongside its terms and its robots.txt, as discussed in robots.txt, AI opt-outs and reservation signals. For many projects, the right response to a heavily protected target is a licence, a partnership or an official API rather than an escalation. Where collection is appropriate, the escalation ladder for protected sites explains the options in order of cost.

Limits of the measurement

  • Homepages only. Deeper pages are often protected differently.
  • One request each, from one network. Results can differ by country, network reputation, time of day and request history; we covered regional differences in which countries get geo-blocked most.
  • No browser. Many challenges are designed to be passed by a real browser running JavaScript; our client could not, by design.
  • Signatures miss things. In-house detection and products that leave no first-request trace are counted as “none detected”, so the true share of protected sites is higher.

The bottom line

A single ordinary request tells you most of what you need to plan a scraping project: who protects the site, whether bot scoring is active, and how a plain client is treated. On the top 1,000 domains, most homepages answered, a quarter sat behind Cloudflare, one in six ran an active bot-management product, and the strictest products challenged or blocked every request we sent.

Run the lookup before writing the scraper, on the pages you need. Let it tell you when to keep things simple, when to plan for a browser, and when to look for an official route to the data instead.

Sources and references

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started