The worst scraper failures do not look like failures. The job runs, the requests succeed, rows arrive in the warehouse on schedule. Then someone in finance asks why the average price tripled overnight, or why half the catalogue is suddenly out of stock, and the investigation leads back to a website redesign three weeks ago that nobody noticed.
This is schema drift: the data keeps flowing, but its shape or meaning has quietly changed. It is the normal failure mode of web data, because the sources change without notice and extractors keep producing output. This guide covers the two layers that catch it: record-level contracts that reject bad values, and batch-level profiles that notice when the data as a whole stops looking like itself. It also shows, with a small test, why you need both.
Key takeaways
- Drift is usually silent. A changed page rarely breaks the extractor; it makes the extractor return something plausible and wrong.
- Record-level contracts check each record against rules: required fields, types, ranges and formats. They catch obvious breakage.
- Batch-level profiles compare each run with a baseline: fill rates, type mix and medians. They catch the breakage contracts cannot see.
- In our test of three realistic drifts, contracts caught one. The other two passed every record-level check and were caught only by the batch profile.
- Alert on drift before data ships, quarantine the batch, and record which site and extractor version produced it.
How drift actually happens
| Cause | What the data does |
|---|---|
| Redesign changes the markup | A selector matches a different element, or nothing |
| Price format changes | ”9.99” becomes “€9.99”, “9,99” or 999 in minor units |
| A field moves behind JavaScript | The static HTML no longer contains it, so it comes back empty |
| Localisation or geo variant | A different currency, language or unit appears for some pages |
| A/B test | A fraction of pages use a new layout, so a fraction of records break |
| Block or challenge page | The page loads, the extractor runs, and returns nothing or rubbish |
The common thread is that none of these raise an exception. The last one, pages that load successfully but are not the page you wanted, is covered in the silent failure rate. The rest are the extractor faithfully processing a page that has changed underneath it.
Layer 1: record-level contracts
A data contract states what a valid record looks like: which fields are required, their types, their allowed ranges and formats. Every record is checked before it is accepted, and violations are counted and quarantined rather than silently stored.
import re
import statistics
from collections import Counter
CONTRACT = {
"name": {"type": str, "required": True},
"price": {"type": float, "required": True, "min": 0.01, "max": 100_000},
"currency": {"type": str, "required": True, "pattern": r"^[A-Z]{3}$"},
"in_stock": {"type": bool, "required": False},
}
def violations(record, contract=CONTRACT):
"""Record-level checks: presence, type, range, format."""
problems = []
for field, rule in contract.items():
value = record.get(field)
if value in (None, ""):
if rule.get("required"):
problems.append(f"{field}: missing")
continue
if rule["type"] is float and isinstance(value, int) and not isinstance(value, bool):
value = float(value)
if not isinstance(value, rule["type"]):
problems.append(f"{field}: expected {rule['type'].__name__}, got {type(value).__name__}")
continue
if "min" in rule and value < rule["min"] or "max" in rule and value > rule["max"]:
problems.append(f"{field}: {value} out of range")
if "pattern" in rule and not re.match(rule["pattern"], value):
problems.append(f"{field}: {value!r} bad format")
return problems
A record like {"name": "x", "price": "€9.99", "currency": "eur"} fails twice: the price is a string, and the currency is not a three-letter uppercase code. Contracts are cheap, explicit and easy to reason about. Their limit is that each record is judged alone, and plenty of drift produces records that are individually valid.
Layer 2: batch-level profiles
A profile summarises a whole batch: for each field, how often it is filled, which types appear and, for numbers, the median. Comparing each run’s profile with a baseline from recent healthy runs shows when the data as a whole changes shape, even if every record passes its contract.
def profile(records, fields=CONTRACT):
"""Batch-level shape: how often each field is filled, its types, and numeric medians."""
n = max(1, len(records))
out = {}
for field in fields:
values = [r.get(field) for r in records]
present = [v for v in values if v not in (None, "")]
numbers = [float(v) for v in present if isinstance(v, (int, float)) and not isinstance(v, bool)]
out[field] = {
"fill_rate": len(present) / n,
"types": Counter(type(v).__name__ for v in present),
"median": statistics.median(numbers) if numbers else None,
"distinct": len(set(map(str, present))),
}
return out
def drift(baseline, current, fill_drop=0.1, median_ratio=3.0):
"""Compare two batch profiles and describe what changed shape."""
alerts = []
for field, base in baseline.items():
cur = current[field]
if base["fill_rate"] - cur["fill_rate"] > fill_drop:
alerts.append(f"{field}: filled {base['fill_rate']:.0%} -> {cur['fill_rate']:.0%}")
if set(cur["types"]) - set(base["types"]):
alerts.append(f"{field}: new types {sorted(set(cur['types']) - set(base['types']))}")
if base["median"] and cur["median"]:
ratio = cur["median"] / base["median"]
if ratio > median_ratio or ratio < 1 / median_ratio:
alerts.append(f"{field}: median {base['median']:g} -> {cur['median']:g}")
return alerts
The thresholds are starting points. A 10-point drop in fill rate or a threefold change in a median is rarely business as usual; tune both per field once you have a few weeks of history.
Why you need both: a small test
We generated a healthy baseline of 1,000 product records, then simulated three realistic breakages and ran both layers on each.
| Drift | What happened | Record-level contract | Batch profile |
|---|---|---|---|
| Formatted price | A redesign put the currency symbol inside the price on 30% of pages | Caught: 300 records rejected | Caught: new type str in price |
| Minor units | Prices started arriving as 6305 instead of 63.05 | Missed: every record passed | Caught: median 63.05 to 6305, and a new type int |
| Broken stock selector | The in-stock field came back empty on 80% of pages | Missed: the field is optional | Caught: filled 100% to 20% |
The contract caught the one drift that produced obviously malformed values. The other two produced records that were individually valid, a positive number within range and an optional field left empty, and only the batch view showed that something had changed. The minor-units case is not hypothetical: a real product endpoint we examined last week returned 10000 for a $100.00 shoe, as described in stop parsing HTML.
What to do when drift fires
- Quarantine the batch. Do not ship data from a batch with drift alerts until someone has looked. A late dataset is better than a wrong one.
- Pinpoint the scope. Break the alert down by site, page type and extractor version. Drift usually starts on one site after one change.
- Compare a sample. Look at a handful of affected records next to the pages they came from. The cause is usually obvious within minutes.
- Fix and backfill. Update the extractor, then re-extract the affected period from stored pages if you keep them, or re-collect if you do not.
- Update the baseline deliberately. When a change is legitimate, such as a site genuinely changing currency, reset the baseline on purpose, never automatically.
Make drift less likely in the first place
Some sources drift less than others. Structured data embedded for search engines changes far less often than page layout, which is why extracting it first reduces how often contracts fire at all. Watching the pages themselves helps too: a template change detected by change detection at scale is an early warning that extraction may be about to break, and a rising rate of contract violations on one site is a strong input to a target health score. Collecting consistently from one market per job also removes a whole class of drift caused by geo and currency variants.
The bottom line
Web data drifts because the web changes without asking. The danger is not that extractors break, but that they keep working on pages that are no longer what they were built for, and produce data that looks fine one record at a time.
Check every record against a contract, and check every batch against its own recent history. In our test, the contract alone caught one drift in three; the batch profile caught all three. Together, they turn “the dashboard looked odd for three weeks” into an alert on the morning it started.
Sources and references
- Test run by Shifter on 29 September 2026 with the code above, on 1,000 generated product records and three simulated drifts.