Knowledge

Schema Drift in Scraped Data: Catching Broken Extractors Before Your Users Do

Scrapers rarely crash when a site changes. They keep running and return subtly wrong data. How data contracts and batch profiling catch drift early.

Matt Brown

Matt Brown

September 29, 2026 · 8 min read

The worst scraper failures do not look like failures. The job runs, the requests succeed, rows arrive in the warehouse on schedule. Then someone in finance asks why the average price tripled overnight, or why half the catalogue is suddenly out of stock, and the investigation leads back to a website redesign three weeks ago that nobody noticed.

This is schema drift: the data keeps flowing, but its shape or meaning has quietly changed. It is the normal failure mode of web data, because the sources change without notice and extractors keep producing output. This guide covers the two layers that catch it: record-level contracts that reject bad values, and batch-level profiles that notice when the data as a whole stops looking like itself. It also shows, with a small test, why you need both.

Key takeaways

  • Drift is usually silent. A changed page rarely breaks the extractor; it makes the extractor return something plausible and wrong.
  • Record-level contracts check each record against rules: required fields, types, ranges and formats. They catch obvious breakage.
  • Batch-level profiles compare each run with a baseline: fill rates, type mix and medians. They catch the breakage contracts cannot see.
  • In our test of three realistic drifts, contracts caught one. The other two passed every record-level check and were caught only by the batch profile.
  • Alert on drift before data ships, quarantine the batch, and record which site and extractor version produced it.

How drift actually happens

CauseWhat the data does
Redesign changes the markupA selector matches a different element, or nothing
Price format changes”9.99” becomes “€9.99”, “9,99” or 999 in minor units
A field moves behind JavaScriptThe static HTML no longer contains it, so it comes back empty
Localisation or geo variantA different currency, language or unit appears for some pages
A/B testA fraction of pages use a new layout, so a fraction of records break
Block or challenge pageThe page loads, the extractor runs, and returns nothing or rubbish

The common thread is that none of these raise an exception. The last one, pages that load successfully but are not the page you wanted, is covered in the silent failure rate. The rest are the extractor faithfully processing a page that has changed underneath it.

Layer 1: record-level contracts

A data contract states what a valid record looks like: which fields are required, their types, their allowed ranges and formats. Every record is checked before it is accepted, and violations are counted and quarantined rather than silently stored.

import re
import statistics
from collections import Counter

CONTRACT = {
    "name":     {"type": str, "required": True},
    "price":    {"type": float, "required": True, "min": 0.01, "max": 100_000},
    "currency": {"type": str, "required": True, "pattern": r"^[A-Z]{3}$"},
    "in_stock": {"type": bool, "required": False},
}


def violations(record, contract=CONTRACT):
    """Record-level checks: presence, type, range, format."""
    problems = []
    for field, rule in contract.items():
        value = record.get(field)
        if value in (None, ""):
            if rule.get("required"):
                problems.append(f"{field}: missing")
            continue
        if rule["type"] is float and isinstance(value, int) and not isinstance(value, bool):
            value = float(value)
        if not isinstance(value, rule["type"]):
            problems.append(f"{field}: expected {rule['type'].__name__}, got {type(value).__name__}")
            continue
        if "min" in rule and value < rule["min"] or "max" in rule and value > rule["max"]:
            problems.append(f"{field}: {value} out of range")
        if "pattern" in rule and not re.match(rule["pattern"], value):
            problems.append(f"{field}: {value!r} bad format")
    return problems

A record like {"name": "x", "price": "€9.99", "currency": "eur"} fails twice: the price is a string, and the currency is not a three-letter uppercase code. Contracts are cheap, explicit and easy to reason about. Their limit is that each record is judged alone, and plenty of drift produces records that are individually valid.

Layer 2: batch-level profiles

A profile summarises a whole batch: for each field, how often it is filled, which types appear and, for numbers, the median. Comparing each run’s profile with a baseline from recent healthy runs shows when the data as a whole changes shape, even if every record passes its contract.

def profile(records, fields=CONTRACT):
    """Batch-level shape: how often each field is filled, its types, and numeric medians."""
    n = max(1, len(records))
    out = {}
    for field in fields:
        values = [r.get(field) for r in records]
        present = [v for v in values if v not in (None, "")]
        numbers = [float(v) for v in present if isinstance(v, (int, float)) and not isinstance(v, bool)]
        out[field] = {
            "fill_rate": len(present) / n,
            "types": Counter(type(v).__name__ for v in present),
            "median": statistics.median(numbers) if numbers else None,
            "distinct": len(set(map(str, present))),
        }
    return out


def drift(baseline, current, fill_drop=0.1, median_ratio=3.0):
    """Compare two batch profiles and describe what changed shape."""
    alerts = []
    for field, base in baseline.items():
        cur = current[field]
        if base["fill_rate"] - cur["fill_rate"] > fill_drop:
            alerts.append(f"{field}: filled {base['fill_rate']:.0%} -> {cur['fill_rate']:.0%}")
        if set(cur["types"]) - set(base["types"]):
            alerts.append(f"{field}: new types {sorted(set(cur['types']) - set(base['types']))}")
        if base["median"] and cur["median"]:
            ratio = cur["median"] / base["median"]
            if ratio > median_ratio or ratio < 1 / median_ratio:
                alerts.append(f"{field}: median {base['median']:g} -> {cur['median']:g}")
    return alerts

The thresholds are starting points. A 10-point drop in fill rate or a threefold change in a median is rarely business as usual; tune both per field once you have a few weeks of history.

Why you need both: a small test

We generated a healthy baseline of 1,000 product records, then simulated three realistic breakages and ran both layers on each.

DriftWhat happenedRecord-level contractBatch profile
Formatted priceA redesign put the currency symbol inside the price on 30% of pagesCaught: 300 records rejectedCaught: new type str in price
Minor unitsPrices started arriving as 6305 instead of 63.05Missed: every record passedCaught: median 63.05 to 6305, and a new type int
Broken stock selectorThe in-stock field came back empty on 80% of pagesMissed: the field is optionalCaught: filled 100% to 20%

The contract caught the one drift that produced obviously malformed values. The other two produced records that were individually valid, a positive number within range and an optional field left empty, and only the batch view showed that something had changed. The minor-units case is not hypothetical: a real product endpoint we examined last week returned 10000 for a $100.00 shoe, as described in stop parsing HTML.

What to do when drift fires

  1. Quarantine the batch. Do not ship data from a batch with drift alerts until someone has looked. A late dataset is better than a wrong one.
  2. Pinpoint the scope. Break the alert down by site, page type and extractor version. Drift usually starts on one site after one change.
  3. Compare a sample. Look at a handful of affected records next to the pages they came from. The cause is usually obvious within minutes.
  4. Fix and backfill. Update the extractor, then re-extract the affected period from stored pages if you keep them, or re-collect if you do not.
  5. Update the baseline deliberately. When a change is legitimate, such as a site genuinely changing currency, reset the baseline on purpose, never automatically.

Make drift less likely in the first place

Some sources drift less than others. Structured data embedded for search engines changes far less often than page layout, which is why extracting it first reduces how often contracts fire at all. Watching the pages themselves helps too: a template change detected by change detection at scale is an early warning that extraction may be about to break, and a rising rate of contract violations on one site is a strong input to a target health score. Collecting consistently from one market per job also removes a whole class of drift caused by geo and currency variants.

The bottom line

Web data drifts because the web changes without asking. The danger is not that extractors break, but that they keep working on pages that are no longer what they were built for, and produce data that looks fine one record at a time.

Check every record against a contract, and check every batch against its own recent history. In our test, the contract alone caught one drift in three; the batch profile caught all three. Together, they turn “the dashboard looked odd for three weeks” into an alert on the morning it started.

Sources and references

  • Test run by Shifter on 29 September 2026 with the code above, on 1,000 generated product records and three simulated drifts.

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started