Scraping

Keep the Raw Response: Archiving Scraped Pages in WARC for Replay

Parsers improve and extractors break, but a page you did not keep cannot be parsed again. How to store raw responses in WARC, with tested code and real sizes.

Matt Brown

Matt Brown

October 1, 2026 · 9 min read

Most scraping pipelines throw the page away the moment they have extracted from it. The parser reads the HTML, writes a row, and the response is gone. That works until the day the extractor turns out to have been wrong for a week, or a new field is needed for last quarter, or someone asks what the page actually said. At that point the only way back is to fetch everything again, and the pages have changed in the meantime.

Keeping the raw response fixes all three problems, and there is a standard format built for it: WARC, the Web ARChive format used by web archives and by Common Crawl. This guide covers what to store, the code to store it, and what it cost on a real test.

Key takeaways

  • Store every response as received, before parsing. Re-extracting from stored responses is cheap; recollecting history is impossible.
  • WARC is the standard container: an open format with request, response and metadata records, readable by many existing tools.
  • Request compressed responses and store them still compressed. On our test of 40 pages, that kept 20.6 MB of HTML down to 3.4 MB on the wire and 3.5 MB on disk.
  • Replaying all 40 pages from the archive took under 0.2 seconds, against 6 to 10 seconds to fetch them.
  • Record how each file was collected, including the exit country, so archived pages can be compared fairly later.

Why keep the raw response

Three situations come up in almost every long-running collection project.

Extractors break silently. A site changes its markup, the parser keeps running, and a field quietly goes empty or wrong. Schema drift monitoring catches it, but only after some bad rows have been written. With raw responses on disk, you fix the extractor and re-run it over the affected days. Without them, those days are lost.

Questions change. Six months in, someone needs a field nobody extracted: a shipping note, a seller rating, a badge. If the pages are archived, it is a batch job. If not, it starts today and has no history.

Evidence needs the original. When collected data supports a decision, a complaint or a dispute, the page as served carries more weight than a row derived from it, which is why court-ready evidence starts with preserved captures.

It also changes how you can work with extraction. Trying a new approach, such as the model-based extraction compared with selectors, becomes an offline experiment on stored pages rather than a new crawl.

What WARC is

WARC is a container format for web captures, maintained by the International Internet Preservation Consortium and standardised as ISO 28500. A WARC file is a sequence of records, each with a short block of text headers followed by content. The record types that matter for scraping are:

Record typeWhat it holds
warcinfoHow the file was made: software, operator, collection notes
requestThe HTTP request as sent, including headers such as Accept-Language
responseThe HTTP response as received: status line, headers and body
metadataAnything else about a capture, linked to it by record ID
revisitA pointer to an earlier identical capture, used to avoid storing duplicates

Each record carries a digest of its content, so a file can be checked for corruption. The specification recommends compressing each record separately with gzip, which keeps files small while still allowing a reader to jump straight to one record. Common Crawl, for example, distributes its crawl data as WARC files.

The practical advantage over a homemade format is tooling. Files written to the standard can be indexed, validated, replayed and searched with existing open-source tools, by people who never saw your code.

The code

The Python library warcio, from the Webrecorder project, reads and writes WARC. The function below fetches a page with requests, writes the request and the response to an open WARC file, and keeps the body exactly as the server sent it:

from io import BytesIO
from urllib.parse import urlsplit

import requests
from warcio.archiveiterator import ArchiveIterator
from warcio.statusandheaders import StatusAndHeaders
from warcio.warcwriter import WARCWriter


def fetch_and_archive(session, url, writer, **kwargs):
    """Fetch url and write the request and the response, byte for byte as received, to a WARC file."""
    response = session.get(url, stream=True, **kwargs)
    # Read the body undecoded, so gzip or br responses are stored exactly as the server sent them.
    raw = response.raw.read(decode_content=False)

    sent = response.request
    path = urlsplit(sent.url)
    request_line = f"{sent.method} {path.path or '/'}{'?' + path.query if path.query else ''} HTTP/1.1"
    request_headers = [("Host", path.netloc)] + list(sent.headers.items())
    request = writer.create_warc_record(
        sent.url, "request", payload=BytesIO(b""),
        http_headers=StatusAndHeaders(request_line, request_headers, is_http_request=True))

    # urllib3 has already removed any chunked framing, so drop the header that describes it.
    headers = [(k, v) for k, v in response.raw.headers.items() if k.lower() != "transfer-encoding"]
    status = f"{response.status_code} {response.reason}"
    record = writer.create_warc_record(
        response.url, "response", payload=BytesIO(raw),
        http_headers=StatusAndHeaders(status, headers, protocol="HTTP/1.1"))

    request.rec_headers.add_header("WARC-Concurrent-To", record.rec_headers.get_header("WARC-Record-ID"))
    writer.write_record(request)
    writer.write_record(record)
    return response.status_code, raw


def replay(path):
    """Yield (url, status, decoded body) for every response in a WARC file, with no network access."""
    with open(path, "rb") as stream:
        for record in ArchiveIterator(stream):
            if record.rec_type == "response":
                status = int(record.http_headers.get_statuscode())
                yield record.rec_headers.get_header("WARC-Target-URI"), status, record.content_stream().read()

And using it through a proxy, with a warcinfo record that says where the pages were collected from:

import os

import requests
from warcio.warcwriter import WARCWriter

from archive import fetch_and_archive, replay

proxy = f"http://{os.environ['SHIFTER_PROXY_USER']}-country-de:{os.environ['SHIFTER_PROXY_PASS']}@p.shifter.io:443"
session = requests.Session()
session.proxies = {"http": proxy, "https": proxy}
session.headers["Accept-Encoding"] = "gzip"
# Identify your collector; some sites refuse the library's default User-Agent.
session.headers["User-Agent"] = "ExampleArchiver/1.0 (+https://example.com/bot)"

urls = ["https://en.wikipedia.org/wiki/Web_archiving", "https://en.wikipedia.org/wiki/Web_crawler"]

with open("crawl-2026-10-01.warc.gz", "wb") as output:
    writer = WARCWriter(output, gzip=True)
    # One warcinfo record per file says how and from where the pages were collected.
    writer.write_record(writer.create_warcinfo_record(
        "crawl-2026-10-01.warc.gz", {"software": "warcio", "description": "exit country: de"}))
    for url in urls:
        fetch_and_archive(session, url, writer, timeout=30)

# Months later, with no network access: re-run a new extractor over the same bytes.
for url, status, body in replay("crawl-2026-10-01.warc.gz"):
    print(status, url, len(body))

Three details in that code came out of testing it.

  • Read the body undecoded. requests normally decompresses responses for you. Reading the raw stream keeps the bytes as sent, which is both more faithful and far smaller. warcio decompresses them again on replay.
  • Drop the chunked header. A quarter of our responses arrived with chunked transfer encoding, which the HTTP library had already unwrapped. Storing the header with the unwrapped body leaves a record that describes itself wrongly. warcio happened to cope; other readers may not.
  • Send a real User-Agent. Our first run of the example was refused with HTTP 403, because the site rejects the HTTP library’s default User-Agent. Identify your collector honestly.

One more gotcha, found in the library’s source rather than in testing: if you accept Brotli-compressed responses (br), install the brotli package. Without it, warcio cannot decode those bodies on replay and returns the still-compressed bytes without raising an error.

What it cost

We archived 40 English Wikipedia articles on 1 October 2026 through a residential exit in Germany, twice: once accepting gzip-compressed responses, and once asking for uncompressed ones.

Compressed responsesUncompressed responses
Pages4040
Transferred3.43 MB20.6 MB
HTML once decoded20.6 MB20.6 MB
WARC file on disk3.54 MB3.51 MB
Time to fetch6 to 10 s9 to 12 s
Time to replay all 400.19 s0.18 s

Every replayed page was byte-identical to what was fetched, checked by SHA-256 hash, and warcio check validated the digests of all 81 records in the file.

Two things stand out. First, storage is cheap: the archive was about 3% larger than the compressed bytes transferred, the difference being the request records and headers. Second, storage was the same either way, because the WARC file is compressed regardless; what changed was bandwidth. Asking for uncompressed responses cost six times as much transfer for an identical archive. On bandwidth-billed collection, that is the difference that shows up on the invoice, as cutting proxy bandwidth costs explains in more detail.

Practical rules

  • Archive before parsing. Write the WARC record first, then extract. If the extractor crashes, the page is still kept.
  • Roll files by size or time. The WARC specification recommends 1 GB as a practical target size per file. Name files with the date and collection, and never append to a file another process is writing.
  • Record the vantage point. A page fetched from Germany and one fetched from the United States can differ in language, price and content. Put the exit country and collection settings in the warcinfo record, as argued in the case for recording where data was observed.
  • Index what you keep. A small index of URL, date, file and offset lets you pull one capture out of terabytes without reading everything. warcio’s command-line index tool produces one.
  • Set a retention policy. Raw pages can contain personal data. Decide how long you keep them, restrict who can read them, and delete on schedule.
  • Skip duplicates deliberately. When a page has not changed since the last capture, a revisit record can point to the earlier copy instead of storing it again, which pairs well with change detection.

The bottom line

Parsing is the part of a scraping pipeline most likely to be wrong and most likely to change, so it should not be the only record of what was collected. Store the raw response in WARC before you parse it, request compressed responses and keep them compressed, and write down where each file was collected from.

On our test, that cost roughly 3.5 MB of disk for 40 pages and turned a crawl that took seconds into a replay that took a fraction of one. The next time an extractor breaks, or a new question arrives about last month, the answer is a batch job rather than a lost week.

Sources and references

  • International Internet Preservation Consortium, The WARC Format 1.1.
  • Webrecorder, warcio, version 1.8.1, used for the code and test above.
  • Common Crawl, Get started, on its use of the WARC format.
  • Test archive of 40 pages collected by Shifter on 1 October 2026, using the code above.

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started