Collect from more than one place, or from one place more than once, and duplicates follow. The same product appears on three marketplaces with three slightly different names. The same company is “Acme Widgets Ltd” in one registry and “ACME WIDGETS LIMITED” in another. The same listing turns up twice because pagination shifted while you were crawling, or because one link carried a tracking parameter and the other did not.
Left alone, duplicates inflate counts, split history across records and quietly corrupt every metric built on top. This guide covers how to resolve them: normalising identifiers, matching on strong keys first, fuzzy matching safely at scale, clustering, and choosing which version of a record survives.
Key takeaways
- Most duplicates are caught by normalising identifiers before any fuzzy matching: canonical URLs, validated product codes and cleaned company names.
- Match on strong identifiers first. A valid barcode or a registry number beats any amount of name similarity.
- Never compare every record with every other. A million records make about 500 billion pairs; blocking reduces that to something tractable.
- Decide which error hurts more. For some jobs a false merge is worse than a missed duplicate; for others it is the reverse.
- Keep provenance. The merged record should still know every source it came from.
Three kinds of duplicate
| Kind | Example | How it is caught |
|---|---|---|
| Repeat collection | The same URL fetched twice, or overlapping pages of results | Canonical URL or source identifier |
| Same entity, different source | One product listed on three marketplaces | Shared identifiers, then fuzzy matching |
| Near-duplicate content | The same article syndicated with small edits | Content similarity, as covered in change detection at scale |
This guide focuses on the first two, where the goal is one record per real-world product, company or listing.
Step 1: normalise identifiers
Normalisation is cheap and catches more duplicates than any clever matching.
URLs. The same page arrives under many URLs. Lowercase the host, drop www., remove tracking parameters such as utm_*, gclid and fbclid, sort the remaining query parameters, and strip trailing slashes. http://www.Shop.com/p/123/?utm_source=x&b=2&a=1 and https://shop.com/p/123?a=1&b=2 become the same key.
Product codes. Global Trade Item Numbers (the numbers behind UPC and EAN barcodes) carry a check digit, which lets you reject mistyped or scraped-wrong codes before trusting them. GS1’s method weights the digits 3, 1, 3, 1 and so on, starting from the digit next to the check digit, sums them, and takes the amount needed to round up to the next multiple of ten. In GS1’s own worked example, the 11-digit body 61414121022 gets a check digit of 0. Pad valid codes to 14 digits so that a 13-digit and a 14-digit version of the same code compare equal.
Company names. Fold case and accents, strip punctuation, and remove legal suffixes such as Ltd, Limited, Inc, GmbH and SA before comparing. “Acme Widgets Ltd.” and “ACME WIDGETS LIMITED” both become acme widgets. Where a company has a registry number or a Legal Entity Identifier, use that instead of the name; resolving companies to identifiers is covered in monitoring your suppliers’ public footprint.
Step 2: match on strong keys first
Once identifiers are normalised, exact matches on them are both fast and reliable. Two records with the same valid barcode are the same product. Two records with the same canonical URL are the same page. Two companies with the same registry number are the same company, whatever they are called.
Structured data helps here too. Many product pages publish barcodes and SKUs in JSON-LD, which is far more reliable than scraping them from the visible page; see stop parsing HTML.
Step 3: fuzzy match, but only within blocks
Records without shared identifiers need fuzzy matching on names, addresses or descriptions. The trap is scale. Comparing every record with every other grows with the square of the dataset: a million records produce roughly 500 billion pairs.
Blocking solves this. Group records by a cheap key that true matches almost always share, such as country plus the first word of the normalised name, or brand plus product category, and compare only within each group. A good blocking key cuts comparisons by orders of magnitude while losing few true matches. Check what it loses by sampling pairs across blocks from time to time.
Within a block, a string similarity score and a threshold decide matches. Start strict, around 0.9, and loosen only after reviewing what the looser setting would merge.
Step 4: cluster, carefully
Matches are pairwise; entities are groups. A union-find structure turns pairs into clusters efficiently. It also carries a risk: transitivity. If A matches B and B matches C, A and C end up together even if they share nothing. Long chains of weak matches are how two different companies get merged. Watch for unusually large clusters and review them before accepting.
The whole pipeline fits in a small module:
import re
import unicodedata
from collections import defaultdict
from difflib import SequenceMatcher
from urllib.parse import urlsplit, urlunsplit, parse_qsl, urlencode
LEGAL_SUFFIXES = {"ltd", "limited", "inc", "incorporated", "llc", "gmbh", "ag", "sa", "sas",
"srl", "bv", "nv", "plc", "co", "corp", "corporation", "company", "oy", "ab"}
TRACKING = re.compile(r"^(utm_|gclid$|fbclid$|mc_|ref$|ref_)")
def gtin_valid(code):
"""GS1 check digit: weights 3,1,3,... from the digit next to the check digit."""
digits = re.sub(r"\D", "", str(code or ""))
if len(digits) not in (8, 12, 13, 14):
return False
body, check = digits[:-1], int(digits[-1])
total = sum(int(d) * (3 if i % 2 == 0 else 1) for i, d in enumerate(reversed(body)))
return (10 - total % 10) % 10 == check
def norm_name(name):
text = unicodedata.normalize("NFKD", name or "").encode("ascii", "ignore").decode().lower()
tokens = [t for t in re.findall(r"[a-z0-9]+", text) if t not in LEGAL_SUFFIXES]
return " ".join(tokens)
def norm_url(url):
parts = urlsplit((url or "").strip())
query = urlencode(sorted((k, v) for k, v in parse_qsl(parts.query) if not TRACKING.match(k.lower())))
host = parts.netloc.lower().removeprefix("www.")
return urlunsplit(("https", host, parts.path.rstrip("/") or "/", query, ""))
def similar(a, b):
return SequenceMatcher(None, a, b).ratio()
def cluster(records, threshold=0.9):
"""Group records that refer to the same entity. Returns lists of record indexes."""
parent = list(range(len(records)))
def find(i):
while parent[i] != i:
parent[i] = parent[parent[i]]
i = parent[i]
return i
def union(i, j):
parent[find(i)] = find(j)
# 1. Exact matches on strong identifiers.
by_key = defaultdict(list)
for i, r in enumerate(records):
if gtin_valid(r.get("gtin")):
by_key["gtin:" + re.sub(r"\D", "", r["gtin"]).zfill(14)].append(i)
if r.get("url"):
by_key["url:" + norm_url(r["url"])].append(i)
for ids in by_key.values():
for j in ids[1:]:
union(ids[0], j)
# 2. Fuzzy name match, only within a cheap blocking key.
blocks = defaultdict(list)
for i, r in enumerate(records):
name = norm_name(r.get("name"))
if name:
blocks[(r.get("country") or "", name.split()[0])].append((i, name))
for members in blocks.values():
for a in range(len(members)):
for b in range(a + 1, len(members)):
if similar(members[a][1], members[b][1]) >= threshold:
union(members[a][0], members[b][0])
groups = defaultdict(list)
for i in range(len(records)):
groups[find(i)].append(i)
return list(groups.values())
On a small test set it merges “Acme Widgets Ltd”, “ACME WIDGETS LIMITED” and a third record sharing the company’s canonical URL, merges two shoe listings whose barcodes differ only by a leading zero, and deliberately leaves “Acme Widget Co” in the United States as a separate entity, because the blocking key includes the country. Whether that last decision is right depends on your data, which is the point of the next step.
Step 5: decide which record survives
A cluster is several versions of one entity, and you need one. Common survivorship rules:
- Most complete wins, field by field: take the non-empty value with the best source for each field rather than one whole record.
- Most recent wins for values that change, such as price and availability.
- Most trusted source wins for values like official names and addresses, for instance a registry over a directory.
Whatever you choose, keep every source identifier and URL on the merged record, with the time each was observed. When a merge turns out to be wrong, provenance is what lets you split it again.
Measure precision and recall
Entity resolution has two error types, and which one matters depends on the job.
| Error | What happens | Worst for |
|---|---|---|
| False merge | Two real entities become one | Company and people-adjacent data, compliance, anything legal |
| Missed duplicate | One entity stays as several records | Counts, market sizing, price comparison |
Label a few hundred candidate pairs by hand, measure both rates, and tune thresholds and blocking keys against the error you care about. A B2B lead database, like the one in building a B2B lead database, usually tolerates a missed duplicate far better than two companies wrongly fused. A price comparison feed, like a real-time competitive price feed, is the other way round: a missed duplicate means a product shows up twice with two prices.
Dedupe early, and at the source
Duplicates cost money before they cost accuracy. Every repeat fetch of the same canonical URL is bandwidth or credits spent for nothing, which is why canonicalising URLs belongs in the crawler’s frontier, not only in the warehouse. The effect on unit cost is covered in cost per clean record.
The bottom line
Deduplication is mostly normalisation. Canonical URLs, validated product codes and cleaned company names catch the bulk of duplicates before any fuzzy matching runs. After that, match on strong keys, fuzzy match only within blocks, cluster with an eye on long chains, keep provenance on every merged record, and measure the errors that matter for your use.
Done well, one real-world thing becomes one record, with its full history attached. Done badly, the dataset looks bigger and less trustworthy at the same time.
Sources and references
- GS1 US, How to calculate a check digit manually.
- GLEIF, Open LEI data.