As of October 7, 2026, Shifter has the largest pool of live US IPs, more than Oxylabs and Bright Data, at a fraction of the cost.Shifter now has the largest pool of live US IPs

See benchmarks

Scraping

Payload Efficiency: What Share of a Scraped Page Is Boilerplate You Paid For

We measured 278 content pages from the top 1,000 sites. The text you want is about 1% of the HTML and a tiny share of what a browser downloads.

Elena Petrova

October 8, 2026 · 11 min read

When you pay for proxy traffic by the gigabyte, every byte of a page costs the same: the paragraph you came for, the menu around it, the analytics script, the hero image and the font file. Most teams know that pages are heavy. Few know how much of what they download is the data they actually keep.

We measured it. On 8 October 2026 we fetched the homepages of the top 1,000 domains in the Tranco ranking, followed one article-style link from each, and measured where the bytes go: on the wire, in the HTML, and in everything a browser downloads. This study shares the results, the code we used, and what they mean for scraping budgets.

Key takeaways

  • The main content of a typical page is about 1% of its HTML. On the median content page, 3.2 KB of main text sat inside 260 KiB of HTML.
  • Compression does most of the saving for free. The same HTML arrived as 46 KiB on the wire, about 5.4 times smaller, so the main text was about 6% of the bytes actually transferred.
  • Inline scripts are the biggest single share of HTML. Across all content pages, inline JavaScript made up 38% of the HTML bytes, inline CSS 15% and inline SVG 10%. Main text was 1.3%.
  • A browser multiplies the bill. The median content page pulled 2.6 MB and 81 requests in a headless browser. The HTML document itself was about 3% of that.
  • Blocking images, media and fonts halves browser traffic. Across all pages it removed 48% of the bytes; on the median page, a lean load was 71% of a full one.

How we measured

  • Sample. The top 1,000 domains of the Tranco list (list Q2K34). Many are API, CDN and tracking hosts with no website, so 431 distinct sites returned a usable homepage. From each homepage we followed the first same-site link that looked like a content page (a path ending in a slug of four or more words, excluding login, legal and help pages), which gave 278 usable content pages.
  • HTML layer. One request per page with a regular desktop browser User-Agent and compression enabled. We recorded the bytes received, decompressed them, and split the HTML into inline scripts, inline styles (including style attributes), inline SVG and JSON-LD. We extracted the main content with trafilatura, an open-source extraction library, and counted its bytes as normalised text.
  • Browser layer. We loaded each content page in headless Chromium, waited for the load event plus three seconds, and summed the bytes transferred for every response by resource type. Then we loaded it again with images, media and fonts blocked.
  • Exclusions. Error responses, non-HTML responses and block pages (challenge pages, “access denied” pages and near-empty responses) were excluded before analysis, so the results describe pages that actually served content.

The HTML layer

Measure (median per page)Content pagesHomepages
Pages measured278431
Bytes on the wire46 KiB51 KiB
HTML after decompression260 KiB270 KiB
Main content text3.2 KB1.4 KB
Main content as a share of the HTML1.2%0.6%
Main content as a share of bytes transferred6.2%3.1%
All visible text as a share of the HTML3.6%2.4%

Content pages carry more text than homepages, as expected, but the overall shape is the same: the text a scraper keeps is a sliver of the document it downloads. Even counting every visible word on the page, including menus, footers and cookie notices, text was under 4% of the HTML.

Where does the rest go? Across all 278 content pages:

Part of the HTMLShare of HTML bytes
Inline JavaScript37.5%
Inline CSS and style attributes15.4%
Inline SVG9.8%
JSON-LD structured data0.9%
Main content text1.3%
Markup, attributes and everything else35.1%

Inline JavaScript is the largest single block: framework state, configuration and tracking code embedded straight into the page. Modern sites often ship their entire page data as a JSON blob in a script tag and render it on the client, which is why scripts outweigh the text they eventually display.

That is also an opportunity. The JSON-LD on these pages averaged under 1% of the HTML, and where a page embeds its data as JSON, that blob is usually far cleaner to parse than the rendered markup. Our guides to extracting JSON-LD instead of parsing HTML and finding the API behind the page cover both.

Compression matters more than anything else at this layer. The median page shrank 5.4 times in transit. Only 10 of the 278 content pages were served uncompressed, but a scraper that does not ask for compression gets the full 260 KiB every time. If your HTTP client sends Accept-Encoding: gzip, deflate, br, you are already paying for 46 KiB, not 260.

The browser layer

Rendering a page in a browser downloads everything the page asks for, not just the document.

Measure (per content page)Value
Pages measured in the browser274
Median bytes transferred, full load2.6 MB
Middle half of pages1.2 to 4.3 MB
Median requests, full load81
HTML document as a share of the full load (median)3%
Main content text as a share of the full load (median)0.13%

By resource type, across all full loads:

Resource typeShare of bytes
Images38.1%
Scripts35.9%
Media (video and audio)7.4%
Fonts6.6%
Stylesheets2.8%
HTML documents2.4%
API calls (fetch and XHR)4.9%
Other1.9%

Blocking images, media and fonts, which a scraper almost never needs, brought the median page down to 1.5 MB and 55 requests. Across all pages, the blocked load transferred 52% of the bytes of the full one. Scripts are harder to block safely, because the page often needs them to render the content you came for.

What it means for a scraping budget

Take a job that collects one million content pages a month and keeps only their main text. Using the medians above:

How the pages are fetchedTraffic per million pages
Main text alone, if you could fetch only thatabout 3.2 GB
HTML only, compressedabout 47 GB
HTML only, uncompressedabout 265 GB
Browser, images, media and fonts blockedabout 1.5 TB
Browser, full loadabout 2.6 TB

The gap between the first row and the last is roughly 800 times. The practical order of savings follows from it:

  1. Fetch HTML without a browser whenever the data is in the HTML. That alone is the difference between gigabytes and terabytes.
  2. Always request compression. It cuts HTML traffic by about five times at no cost.
  3. When you must render, block images, media and fonts. It roughly halves browser traffic.
  4. Look for a smaller source of the same data: JSON-LD, an embedded JSON blob or the API the page itself calls.

Our guide to cutting proxy bandwidth costs covers each of these techniques in practice, and estimating your monthly bandwidth shows how to turn page weights into a plan.

Measure your own pages

Medians across the top sites are a starting point; your targets are what matter. The function below fetches one page and reports where its bytes go, using the same method as the study:

import gzip
import re
import zlib

import brotli
import requests
import trafilatura
from lxml import html as lxml_html


def decode(raw, content_encoding):
    """Undo Content-Encoding by hand, so we can count the compressed bytes first."""
    for coding in reversed([c.strip() for c in content_encoding.lower().split(",") if c.strip()]):
        if coding == "gzip":
            raw = gzip.decompress(raw)
        elif coding == "br":
            raw = brotli.decompress(raw)
        elif coding == "deflate":
            raw = zlib.decompress(raw)
    return raw


def size(text):
    return len(text.encode("utf-8"))


def payload_breakdown(url, session=None):
    """Where the bytes of one HTML page go, from the wire down to the main content."""
    session = session or requests.Session()
    response = session.get(url, timeout=30, stream=True,
                           headers={"Accept-Encoding": "gzip, deflate, br"})
    wire = response.raw.read(decode_content=False)
    page = decode(wire, response.headers.get("Content-Encoding", "")).decode(
        response.encoding or "utf-8", errors="replace")
    doc = lxml_html.document_fromstring(page)

    scripts = doc.xpath("//script")
    json_ld = sum(size(s.text or "") for s in scripts if s.get("type") == "application/ld+json")
    inline_js = sum(size(s.text or "") for s in scripts
                    if not s.get("src") and s.get("type") != "application/ld+json")
    inline_css = sum(size(s.text or "") for s in doc.xpath("//style")) + sum(size(v) for v in doc.xpath("//@style"))
    inline_svg = sum(size(lxml_html.tostring(s, encoding="unicode"))
                     for s in doc.xpath("//*[local-name()='svg'][not(ancestor::*[local-name()='svg'])]"))
    main_text = trafilatura.extract(page, include_tables=True) or ""

    return {
        "status": response.status_code,
        "wire_bytes": len(wire),
        "html_bytes": size(page),
        "inline_js": inline_js,
        "inline_css": inline_css,
        "inline_svg": inline_svg,
        "json_ld": json_ld,
        "main_text": size(re.sub(r"\s+", " ", main_text).strip()),
    }

It reads the response body before decompression, so wire_bytes is what you actually transferred, then decodes it by hand. Run it on a few pages from each of your targets:

from payload import payload_breakdown

import requests

session = requests.Session()
session.headers["User-Agent"] = "ExampleStudy/1.0 (+https://example.com/bot)"
b = payload_breakdown("https://en.wikipedia.org/wiki/Web_scraping", session)
print(b)
print(f"main content: {b['main_text'] / b['html_bytes']:.1%} of the HTML, "
      f"{b['main_text'] / b['wire_bytes']:.1%} of the bytes transferred")
{'status': 200, 'wire_bytes': 46860, 'html_bytes': 236286, 'inline_js': 6964, 'inline_css': 6303, 'inline_svg': 0, 'json_ld': 640, 'main_text': 26979}
main content: 11.4% of the HTML, 57.6% of the bytes transferred

Wikipedia is an efficient page by these standards: its main text is over 11% of the HTML, nearly ten times the median we measured. Your targets will fall somewhere on that range, and knowing where tells you whether HTML-only collection, compression or a different data source will save the most.

Limits of the measurement

  • Top sites only. The top 1,000 domains are large, well-engineered sites. Smaller sites may be lighter or much heavier.
  • One content page per site. We followed the first article-style link on each homepage. Product pages, search results and listing pages may differ.
  • Main content is an estimate. Extraction libraries can miss or over-include content, particularly on pages that are mostly navigation. Treat the main-text figures as approximate; the byte measurements are exact.
  • One snapshot, one network. Pages were fetched once, on 8 October 2026, from a single network in Europe. Sites serve different pages, ads and media by location and over time.
  • Browser loads were capped. We stopped measuring three seconds after the load event. Pages that keep loading content afterwards would transfer more than we recorded.

FAQ

What share of a web page is the actual content?

On the median content page among the top 1,000 sites, the main text was about 1.2% of the HTML and about 6% of the compressed bytes transferred. In a full browser load, it was about 0.13% of the bytes.

Does blocking images reduce proxy bandwidth?

Yes. Blocking images, media and fonts in a headless browser removed 48% of the bytes across our sample. Images alone were 38% of browser traffic.

Is it cheaper to scrape without a browser?

Usually by a large margin. The median content page was 46 KiB as compressed HTML and 2.6 MB in a full browser load, a difference of more than 50 times.

Should I request compressed responses when scraping?

Yes. The median page was 5.4 times smaller compressed. Most HTTP clients request compression by default, but check, because a client that does not pays for the full size of every page.

The bottom line

The data a scraper keeps is a tiny share of what it downloads: about 1% of a typical page’s HTML and a fraction of a percent of what a browser pulls. Most of that overhead is avoidable. Fetch HTML instead of rendering when you can, keep compression on, block heavy resources when you must render, and look for structured data or APIs that carry the same information in far fewer bytes.

Sources and references

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started