Scraping

Change Detection at Scale: Diffing Web Pages Without Drowning in Noise

Pages change on every load: ads, timestamps, tokens. How to tell real changes from noise with normalisation, field-level diffs and simhash, with tested code.

Chris Collins

Chris Collins

September 27, 2026 · 9 min read

Every monitoring job eventually asks the same question: did this page change? It sounds like a hash comparison. It is not. Modern pages change on every load, with rotating ads, timestamps, recommendation blocks, session tokens and live counters, so a naive comparison reports a change every time and the alert channel fills with noise until everyone mutes it. The opposite failure is quieter and worse: a pipeline so tuned to ignore noise that it misses the price cut or the removed certificate it was built to catch.

This guide covers how to detect real changes at scale: comparing fields rather than pages, normalising what you cannot avoid comparing, measuring how much a page changed rather than whether it changed at all, and routing each kind of change to the right place.

Key takeaways

  • Raw byte comparison is useless for most sites. We fetched a news homepage twice, 45 seconds apart: the bytes differed, and even the cleaned-up text differed.
  • Compare fields, not pages, wherever you can. A price, a title or a stock status either changed or it did not.
  • Where you must compare whole pages, normalise first, then measure similarity. Simhash, the fingerprinting technique Google described for near-duplicate detection across 8 billion pages, turns “did it change?” into “how much did it change?”
  • Classify every change: field change, content change or noise. Each deserves a different response.
  • Tune per site and per page type. A threshold that is right for a product page is wrong for a news homepage.

Why naive diffing fails

To see the problem concretely, we fetched the international front page of a major news site twice, 45 seconds apart. The raw HTML differed, as expected. More interesting: after stripping scripts, styles and markup, and removing timestamps and tokens, the visible text still differed. Live pages update continuously. A monitor that alerts on any difference would have alerted on this page every single time it ran.

The sources of noise are predictable:

NoiseExampleTypical fix
Embedded scripts and dataAnalytics payloads, JSON state, CSRF tokensStrip script and style blocks before comparing
Time-based text”Updated 3 minutes ago”, clock times, datesRemove with patterns, or compare fields that exclude them
Rotating modulesAds, “trending now”, recommendationsCompare only the region or fields you care about
Personalisation and experimentsA/B test variants, location-specific blocksCollect from a fixed vantage point with consistent settings
OrderingLists that shuffle on each loadSort before comparing

Some noise is really a difference in who asked. A page loaded from Germany can differ from the same page loaded from the United States for reasons that have nothing to do with a change. Keep the vantage point fixed for each monitored page, so the only thing that varies between runs is time. With a residential gateway, that means pinning the country, for example customer-USERNAME-country-de, for every run of that job.

Principle 1: compare fields, not pages

The most reliable change detector does not diff pages at all. It extracts the fields that matter, such as price, availability, title, a certificate list or an address, and compares those. A field changed or it did not, and noise elsewhere on the page is irrelevant.

Structured data makes this easier than it sounds. Many pages publish their key fields as JSON-LD, which is far more stable than the layout around it; extracting it is covered in stop parsing HTML. Where fields come from selectors instead, compare the extracted values, never the HTML around them.

Field comparison also makes alerts useful. “Price changed from 89.00 to 79.00” is actionable. “The page changed” is not.

Principle 2: normalise what you must compare

Some monitoring really is about whole pages: terms of service, a supplier’s “about” page, a sustainability statement, a policy page. For those, strip what never matters before comparing:

  • Remove non-content blocks: scripts, styles, templates, inline SVG.
  • Keep visible text only, with whitespace collapsed and case folded.
  • Remove volatile fragments: clock times, ISO dates, relative times, tokens and long hexadecimal identifiers.
  • Restrict to a region when a page has a stable main content area, such as the article body or the main element.

Normalisation removes a large share of noise. It does not remove all of it, which is why the next step matters.

Principle 3: measure how much, not whether

Instead of asking whether normalised text is identical, ask how similar it is. Simhash is a good tool for this. It reduces a document to a 64-bit fingerprint with a useful property: similar documents get fingerprints that differ in only a few bit positions. Google described using it for near-duplicate detection in “Detecting Near-Duplicates for Web Crawling” (WWW 2007), where the authors validated that “for a repository of 8B web-pages, 64-bit simhash fingerprints and k = 3 are reasonable”, meaning pages whose fingerprints differ in at most three bits could be treated as near-duplicates.

That makes change detection a question of distance. In our test, the two fetches of the news homepage, 45 seconds apart, came out 2 bits apart: below the threshold, so noise. The same homepage compared with an unrelated article came out 34 bits apart: plainly different content.

import hashlib
import re

VOLATILE = [
    r"\b\d{1,2}:\d{2}(?::\d{2})?\s?(?:am|pm|AM|PM)?\b",          # clock times
    r"\b\d{4}-\d{2}-\d{2}(?:T[\d:.]+Z?)?\b",                      # ISO dates
    r"\b\d+\s+(?:second|minute|hour|day)s?\s+ago\b",              # relative times
    r"\b(?:csrf|nonce|token|session|sid)[\w-]*[=:]\s*[\w-]+",      # tokens
    r"\b[0-9a-f]{24,}\b",                                         # long hex ids
]


def normalise(html):
    """Visible text only, with volatile fragments removed."""
    html = re.sub(r"(?is)<(script|style|noscript|svg|template)\b.*?</\1>", " ", html)
    text = re.sub(r"(?s)<[^>]+>", " ", html)
    for pattern in VOLATILE:
        text = re.sub(pattern, " ", text, flags=re.I)
    return re.sub(r"\s+", " ", text).strip().lower()


def simhash(text, bits=64):
    """Charikar-style simhash over word 3-grams."""
    words = text.split()
    grams = [" ".join(words[i:i + 3]) for i in range(max(1, len(words) - 2))]
    weights = [0] * bits
    for gram in grams:
        h = int.from_bytes(hashlib.blake2b(gram.encode(), digest_size=8).digest(), "big")
        for i in range(bits):
            weights[i] += 1 if h >> i & 1 else -1
    return sum(1 << i for i in range(bits) if weights[i] > 0)


def distance(a, b):
    return bin(a ^ b).count("1")


def compare(old_html, new_html, old_fields=None, new_fields=None, threshold=3):
    """Classify a change as none, noise, content or field-level."""
    old_fields, new_fields = old_fields or {}, new_fields or {}
    field_changes = {k: (old_fields.get(k), new_fields.get(k))
                     for k in set(old_fields) | set(new_fields)
                     if old_fields.get(k) != new_fields.get(k)}
    if field_changes:
        return {"kind": "field", "changes": field_changes}
    if old_html == new_html:
        return {"kind": "none"}
    d = distance(simhash(normalise(old_html)), simhash(normalise(new_html)))
    return {"kind": "content" if d > threshold else "noise", "distance": d}

Treat the threshold as a starting point, not a constant. The WWW 2007 figure was chosen for finding duplicates among billions of pages, not for monitoring one page over time. Calibrate it per page type: fetch each monitored page several times in quick succession, where nothing meaningful should change, and set the threshold just above the distances you observe. Short pages need more care, because a few changed words move a small document’s fingerprint further than a long one’s.

Principle 4: classify, then route

A detector that returns only “changed” pushes the real work onto whoever reads the alert. Return a kind, and route by it:

KindMeaningWhere it goes
fieldA value you track changedStraight into the dataset and, if it matters, an alert
contentThe page’s substance changed beyond noiseA human review queue, with a text diff attached
noiseThe page moved, but within normal variationLogged for calibration, never alerted
noneByte-identicalNothing

Two refinements pay for themselves. Keep the previous snapshot and a readable text diff alongside every content change, so a reviewer can judge it in seconds. And record the rate of each kind per site: a site whose noise distance creeps up over weeks is changing its templates, and a site that suddenly produces no changes at all may be serving you a stale or blocked page, a failure mode covered in the silent failure rate.

Scaling it up

Change detection at scale is mostly a storage and scheduling problem.

  • Store fingerprints and fields, not full pages, for most runs. A 64-bit fingerprint and a handful of fields are tiny. Keep full snapshots only when a change is detected, or on a slower schedule for audit.
  • Revisit by expected change. Pages that rarely change do not need hourly checks. Cost-aware crawl scheduling sets out how to spend a fetch budget where change is likely, and every none or noise result is evidence for that model.
  • Use cheap signals first. Where a site sends reliable ETag or Last-Modified headers, a conditional request that returns 304 Not Modified answers the question for almost no bandwidth.
  • Watch the watcher. Canary pages with known change patterns tell you whether the detector itself still works.

Where this is used

The same machinery sits under very different jobs: price and stock monitoring, where field changes are the whole point, as in building a real-time competitive price feed; supplier and compliance monitoring, where whole-page content changes matter, as in monitoring your suppliers’ public footprint and verifying ESG claims; and preserving evidence when a change is legally significant.

The bottom line

“Did this page change?” is the wrong question for the modern web, because the answer is almost always yes. The useful questions are “did a field I care about change?” and, where you must compare whole pages, “did the substance change beyond this page’s normal variation?”

Compare fields wherever you can. Normalise what you cannot avoid comparing, measure similarity rather than equality, calibrate thresholds per page type, and route each kind of change to the place that can act on it. The result is a change feed people trust, which is the only kind that gets read.

Sources and references

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started