When you pay for proxy traffic by the gigabyte, every byte of a page costs the same: the paragraph you came for, the menu around it, the analytics script, the hero image and the font file. Most teams know that pages are heavy. Few know how much of what they download is the data they actually keep.
We measured it. On 8 October 2026 we fetched the homepages of the top 1,000 domains in the Tranco ranking, followed one article-style link from each, and measured where the bytes go: on the wire, in the HTML, and in everything a browser downloads. This study shares the results, the code we used, and what they mean for scraping budgets.
Key takeaways
- The main content of a typical page is about 1% of its HTML. On the median content page, 3.2 KB of main text sat inside 260 KiB of HTML.
- Compression does most of the saving for free. The same HTML arrived as 46 KiB on the wire, about 5.4 times smaller, so the main text was about 6% of the bytes actually transferred.
- Inline scripts are the biggest single share of HTML. Across all content pages, inline JavaScript made up 38% of the HTML bytes, inline CSS 15% and inline SVG 10%. Main text was 1.3%.
- A browser multiplies the bill. The median content page pulled 2.6 MB and 81 requests in a headless browser. The HTML document itself was about 3% of that.
- Blocking images, media and fonts halves browser traffic. Across all pages it removed 48% of the bytes; on the median page, a lean load was 71% of a full one.
How we measured
- Sample. The top 1,000 domains of the Tranco list (list Q2K34). Many are API, CDN and tracking hosts with no website, so 431 distinct sites returned a usable homepage. From each homepage we followed the first same-site link that looked like a content page (a path ending in a slug of four or more words, excluding login, legal and help pages), which gave 278 usable content pages.
- HTML layer. One request per page with a regular desktop browser User-Agent and compression enabled. We recorded the bytes received, decompressed them, and split the HTML into inline scripts, inline styles (including
styleattributes), inline SVG and JSON-LD. We extracted the main content with trafilatura, an open-source extraction library, and counted its bytes as normalised text. - Browser layer. We loaded each content page in headless Chromium, waited for the load event plus three seconds, and summed the bytes transferred for every response by resource type. Then we loaded it again with images, media and fonts blocked.
- Exclusions. Error responses, non-HTML responses and block pages (challenge pages, “access denied” pages and near-empty responses) were excluded before analysis, so the results describe pages that actually served content.
The HTML layer
| Measure (median per page) | Content pages | Homepages |
|---|---|---|
| Pages measured | 278 | 431 |
| Bytes on the wire | 46 KiB | 51 KiB |
| HTML after decompression | 260 KiB | 270 KiB |
| Main content text | 3.2 KB | 1.4 KB |
| Main content as a share of the HTML | 1.2% | 0.6% |
| Main content as a share of bytes transferred | 6.2% | 3.1% |
| All visible text as a share of the HTML | 3.6% | 2.4% |
Content pages carry more text than homepages, as expected, but the overall shape is the same: the text a scraper keeps is a sliver of the document it downloads. Even counting every visible word on the page, including menus, footers and cookie notices, text was under 4% of the HTML.
Where does the rest go? Across all 278 content pages:
| Part of the HTML | Share of HTML bytes |
|---|---|
| Inline JavaScript | 37.5% |
| Inline CSS and style attributes | 15.4% |
| Inline SVG | 9.8% |
| JSON-LD structured data | 0.9% |
| Main content text | 1.3% |
| Markup, attributes and everything else | 35.1% |
Inline JavaScript is the largest single block: framework state, configuration and tracking code embedded straight into the page. Modern sites often ship their entire page data as a JSON blob in a script tag and render it on the client, which is why scripts outweigh the text they eventually display.
That is also an opportunity. The JSON-LD on these pages averaged under 1% of the HTML, and where a page embeds its data as JSON, that blob is usually far cleaner to parse than the rendered markup. Our guides to extracting JSON-LD instead of parsing HTML and finding the API behind the page cover both.
Compression matters more than anything else at this layer. The median page shrank 5.4 times in transit. Only 10 of the 278 content pages were served uncompressed, but a scraper that does not ask for compression gets the full 260 KiB every time. If your HTTP client sends Accept-Encoding: gzip, deflate, br, you are already paying for 46 KiB, not 260.
The browser layer
Rendering a page in a browser downloads everything the page asks for, not just the document.
| Measure (per content page) | Value |
|---|---|
| Pages measured in the browser | 274 |
| Median bytes transferred, full load | 2.6 MB |
| Middle half of pages | 1.2 to 4.3 MB |
| Median requests, full load | 81 |
| HTML document as a share of the full load (median) | 3% |
| Main content text as a share of the full load (median) | 0.13% |
By resource type, across all full loads:
| Resource type | Share of bytes |
|---|---|
| Images | 38.1% |
| Scripts | 35.9% |
| Media (video and audio) | 7.4% |
| Fonts | 6.6% |
| Stylesheets | 2.8% |
| HTML documents | 2.4% |
| API calls (fetch and XHR) | 4.9% |
| Other | 1.9% |
Blocking images, media and fonts, which a scraper almost never needs, brought the median page down to 1.5 MB and 55 requests. Across all pages, the blocked load transferred 52% of the bytes of the full one. Scripts are harder to block safely, because the page often needs them to render the content you came for.
What it means for a scraping budget
Take a job that collects one million content pages a month and keeps only their main text. Using the medians above:
| How the pages are fetched | Traffic per million pages |
|---|---|
| Main text alone, if you could fetch only that | about 3.2 GB |
| HTML only, compressed | about 47 GB |
| HTML only, uncompressed | about 265 GB |
| Browser, images, media and fonts blocked | about 1.5 TB |
| Browser, full load | about 2.6 TB |
The gap between the first row and the last is roughly 800 times. The practical order of savings follows from it:
- Fetch HTML without a browser whenever the data is in the HTML. That alone is the difference between gigabytes and terabytes.
- Always request compression. It cuts HTML traffic by about five times at no cost.
- When you must render, block images, media and fonts. It roughly halves browser traffic.
- Look for a smaller source of the same data: JSON-LD, an embedded JSON blob or the API the page itself calls.
Our guide to cutting proxy bandwidth costs covers each of these techniques in practice, and estimating your monthly bandwidth shows how to turn page weights into a plan.
Measure your own pages
Medians across the top sites are a starting point; your targets are what matter. The function below fetches one page and reports where its bytes go, using the same method as the study:
import gzip
import re
import zlib
import brotli
import requests
import trafilatura
from lxml import html as lxml_html
def decode(raw, content_encoding):
"""Undo Content-Encoding by hand, so we can count the compressed bytes first."""
for coding in reversed([c.strip() for c in content_encoding.lower().split(",") if c.strip()]):
if coding == "gzip":
raw = gzip.decompress(raw)
elif coding == "br":
raw = brotli.decompress(raw)
elif coding == "deflate":
raw = zlib.decompress(raw)
return raw
def size(text):
return len(text.encode("utf-8"))
def payload_breakdown(url, session=None):
"""Where the bytes of one HTML page go, from the wire down to the main content."""
session = session or requests.Session()
response = session.get(url, timeout=30, stream=True,
headers={"Accept-Encoding": "gzip, deflate, br"})
wire = response.raw.read(decode_content=False)
page = decode(wire, response.headers.get("Content-Encoding", "")).decode(
response.encoding or "utf-8", errors="replace")
doc = lxml_html.document_fromstring(page)
scripts = doc.xpath("//script")
json_ld = sum(size(s.text or "") for s in scripts if s.get("type") == "application/ld+json")
inline_js = sum(size(s.text or "") for s in scripts
if not s.get("src") and s.get("type") != "application/ld+json")
inline_css = sum(size(s.text or "") for s in doc.xpath("//style")) + sum(size(v) for v in doc.xpath("//@style"))
inline_svg = sum(size(lxml_html.tostring(s, encoding="unicode"))
for s in doc.xpath("//*[local-name()='svg'][not(ancestor::*[local-name()='svg'])]"))
main_text = trafilatura.extract(page, include_tables=True) or ""
return {
"status": response.status_code,
"wire_bytes": len(wire),
"html_bytes": size(page),
"inline_js": inline_js,
"inline_css": inline_css,
"inline_svg": inline_svg,
"json_ld": json_ld,
"main_text": size(re.sub(r"\s+", " ", main_text).strip()),
}
It reads the response body before decompression, so wire_bytes is what you actually transferred, then decodes it by hand. Run it on a few pages from each of your targets:
from payload import payload_breakdown
import requests
session = requests.Session()
session.headers["User-Agent"] = "ExampleStudy/1.0 (+https://example.com/bot)"
b = payload_breakdown("https://en.wikipedia.org/wiki/Web_scraping", session)
print(b)
print(f"main content: {b['main_text'] / b['html_bytes']:.1%} of the HTML, "
f"{b['main_text'] / b['wire_bytes']:.1%} of the bytes transferred")
{'status': 200, 'wire_bytes': 46860, 'html_bytes': 236286, 'inline_js': 6964, 'inline_css': 6303, 'inline_svg': 0, 'json_ld': 640, 'main_text': 26979}
main content: 11.4% of the HTML, 57.6% of the bytes transferred
Wikipedia is an efficient page by these standards: its main text is over 11% of the HTML, nearly ten times the median we measured. Your targets will fall somewhere on that range, and knowing where tells you whether HTML-only collection, compression or a different data source will save the most.
Limits of the measurement
- Top sites only. The top 1,000 domains are large, well-engineered sites. Smaller sites may be lighter or much heavier.
- One content page per site. We followed the first article-style link on each homepage. Product pages, search results and listing pages may differ.
- Main content is an estimate. Extraction libraries can miss or over-include content, particularly on pages that are mostly navigation. Treat the main-text figures as approximate; the byte measurements are exact.
- One snapshot, one network. Pages were fetched once, on 8 October 2026, from a single network in Europe. Sites serve different pages, ads and media by location and over time.
- Browser loads were capped. We stopped measuring three seconds after the load event. Pages that keep loading content afterwards would transfer more than we recorded.
FAQ
What share of a web page is the actual content?
On the median content page among the top 1,000 sites, the main text was about 1.2% of the HTML and about 6% of the compressed bytes transferred. In a full browser load, it was about 0.13% of the bytes.
Does blocking images reduce proxy bandwidth?
Yes. Blocking images, media and fonts in a headless browser removed 48% of the bytes across our sample. Images alone were 38% of browser traffic.
Is it cheaper to scrape without a browser?
Usually by a large margin. The median content page was 46 KiB as compressed HTML and 2.6 MB in a full browser load, a difference of more than 50 times.
Should I request compressed responses when scraping?
Yes. The median page was 5.4 times smaller compressed. Most HTTP clients request compression by default, but check, because a client that does not pays for the full size of every page.
The bottom line
The data a scraper keeps is a tiny share of what it downloads: about 1% of a typical page’s HTML and a fraction of a percent of what a browser pulls. Most of that overhead is avoidable. Fetch HTML instead of rendering when you can, keep compression on, block heavy resources when you must render, and look for structured data or APIs that carry the same information in far fewer bytes.
Sources and references
- Tranco, list Q2K34, the research-oriented ranking of top sites used for the sample.
- Barbaresi, A. (2021). Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction. ACL 2021 System Demonstrations.
- Pages measured by Shifter on 8 October 2026 using the code above and headless Chromium.