If you compare prices, content or availability across markets, the first problem is finding the same page in every market. Guessing URL patterns works on some sites and fails on most: one site uses /de/, another de.example.com, another a query parameter, and a fourth a separate domain per country.
Many sites already publish the answer. The hreflang annotation, which sites add for search engines, lists every language and country version of a page and where it lives. Read it, and a single fetch gives you the whole set. This guide explains how it works, what we found when we checked how top sites use it, and why holding the right URL is not always enough to see the right page.
Key takeaways
- On 1 October 2026, 82 of the 253 top-site homepages we fetched (32%) declared hreflang alternates, with a median of 17.5 versions each and up to 157.
- One fetch of an annotated page gives you every localized URL for that page, with no guessing of URL patterns.
- The annotations are not always clean: 12 of the 82 sites used at least one code Google does not support, such as
en-uk,es-419,en_GBwith an underscore, or the country before the language. - 86% of the alternates we checked linked straight back, as Google requires. Most of the rest pointed at a different URL for the same page, or redirected.
- Some versions open only from the right place. A news site’s US homepage and a marketplace’s German storefront redirected our connection from Romania elsewhere, and loaded normally through exits in the United States and Germany.
How hreflang works
A page lists its alternates in its <head>:
<link rel="alternate" hreflang="en" href="https://example.com/" />
<link rel="alternate" hreflang="de" href="https://example.com/de/" />
<link rel="alternate" hreflang="pt-BR" href="https://example.com/pt/" />
<link rel="alternate" hreflang="x-default" href="https://example.com/" />
Each value is a language code, optionally followed by a script or a region: de for German anywhere, en-GB for English in the United Kingdom, zh-Hant for Traditional Chinese. x-default names the fallback for visitors who match none of them. Google’s documentation sets the rules that most sites follow:
- Languages use ISO 639-1 codes, scripts ISO 15924 and regions ISO 3166-1 alpha-2. A region on its own is not valid, and codes outside those standards, such as
es-419for Latin American Spanish, are not supported. - URLs must be fully qualified, including
https://. - Each version must list itself and every other version, and if two pages do not point to each other, the annotations may be ignored.
- The same information can be given in the HTML, in an HTTP
Linkheader (useful for PDFs) or in a sitemap.
For data collection, the third rule is the useful one. Because every version is supposed to list every other, any one of them is a complete map of the set.
What we found on top sites
We fetched the homepages of the top 500 domains in the Tranco ranking of popular sites. 253 distinct sites returned an HTML homepage; the rest were content networks, API hosts and other domains without one, or duplicates of a site already counted.
| Measure | Result |
|---|---|
| Homepages with hreflang annotations | 82 of 253 (32%) |
| Alternates per annotated page | median 17.5, maximum 157 |
Pages with an x-default | 63 of 82 |
Homepages sending hreflang in an HTTP Link header | 0 |
| Sites with at least one unsupported code | 12 of 82 |
| Sites using relative URLs | 2 |
| Sites listing the same code twice | 6 |
The unsupported codes fell into a few patterns:
| Pattern | Example | Sites |
|---|---|---|
| Country before language, plus regional labels | mx-es, cz-cs, emea_africa-en | 1 (48 codes) |
| Underscore instead of a hyphen | en_GB, de_DE | 1 (25 codes) |
| Region code that is not ISO 3166-1 | en-uk, en-eu, sq-xk | 3 |
| Latin American Spanish | es-419 | 2 |
| Language code that is not ISO 639-1 | ceb, skr, the outdated iw for Hebrew | 3 |
| Invented codes | tc, zh-FT, ms-en | 3 |
Some sites appear in more than one row. The country-first pattern is the most treacherous for a parser, because several of those codes are accidentally valid in reverse: ca is the code for Catalan, id for Indonesian and th for Thai, so ca-en reads as Catalan rather than English for Canada. Never derive a market from the code alone without checking it makes sense.
Return links
We took up to two alternates from each annotated homepage, 155 in all, fetched them, and checked whether each listed the homepage back.
| Outcome | Alternates |
|---|---|
| Linked straight back | 133 (86%) |
| Linked back to a different URL for the same page | 9 |
| Redirected somewhere else | 9 |
| Genuinely did not list the page | 4, on two sites |
The second row is the subtle one. On several sites, the homepage we landed on lived at one URL while the set named another: / against /en/, /homepage or /home.html, or a bare URL against one with a language parameter. The annotations were consistent; our entry point was simply not the URL the site considers canonical. When you collect, key each version by the URL the set itself uses, not the one you happened to arrive at.
The right URL is not always enough
The redirects were the most instructive part. We fetched each redirecting alternate again, directly and through Shifter exits in the target country, with English and with local Accept-Language headers:
| Site type | Alternate | From Romania | From the target country | Decided by |
|---|---|---|---|---|
| News site | US homepage | Redirected to the international edition | Loaded, via a US exit | Location |
| Marketplace | German storefront | Redirected to the global site | Loaded, via a German exit | Location |
| Video-call vendor | Japanese homepage | Loaded only with Japanese Accept-Language | Same, via a Japanese exit | Language |
| Games publisher | Arabic homepage | Redirected | Redirected, via a Saudi exit | Neither: listed but unreachable |
Two sites decided by the visitor’s location, the marketplace regardless of language, so their own annotations pointed at pages a visitor from elsewhere could not open. One decided by language, regardless of location. And one listed a version that redirected in every combination we tried, which is worth knowing before you build a market comparison on it.
This is where hreflang and exit location meet. The annotations tell you which versions exist and where; collecting them faithfully means requesting each one the way a local visitor would, from that country and with that language, as covered in matching proxy geo, timezone and locale. Otherwise you risk recording the international version under a German label.
The code
The module below reads a page’s HTML alternates, flags unsupported codes and relative URLs, and checks return links, reporting why each failing alternate fails:
from html.parser import HTMLParser
from urllib.parse import urljoin, urlsplit, urlunsplit
import requests
from babel import Locale
# Google accepts ISO 639-1 languages, ISO 15924 scripts and ISO 3166-1 alpha-2 regions.
LANGUAGES = {code for code in Locale("en").languages if len(code) == 2}
NOT_ISO_3166 = {"AC", "CP", "CQ", "DG", "EA", "EU", "EZ", "IC", "QO", "TA", "UN", "XA", "XB", "XK", "ZZ"}
REGIONS = {code for code in Locale("en").territories if code.isalpha()} - NOT_ISO_3166
SCRIPTS = set(Locale("en").scripts)
class AlternateLinks(HTMLParser):
"""Collect <link rel="alternate" hreflang="..."> elements from a page's markup."""
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
a = dict(attrs)
if tag == "link" and "alternate" in (a.get("rel") or "").lower().split() and a.get("hreflang") and a.get("href"):
self.links.append((a["hreflang"].strip(), a["href"].strip()))
def valid_hreflang(code):
"""True for x-default, a language, or language-region / language-script(-region) codes."""
if code.lower() == "x-default":
return True
parts = code.split("-")
if parts[0].lower() not in LANGUAGES:
return False
rest = parts[1:]
if rest and rest[0].title() in SCRIPTS:
rest = rest[1:]
if rest and rest[0].upper() in REGIONS:
rest = rest[1:]
return not rest
def hreflang_map(url, session=None, timeout=20):
"""Fetch a page and return its declared alternates as {hreflang: absolute URL}, plus any problems found."""
session = session or requests.Session()
response = session.get(url, timeout=timeout)
parser = AlternateLinks()
parser.feed(response.text) # link tags are normally ASCII, so the decoding choice rarely matters here
alternates, problems = {}, []
for code, href in parser.links:
if not href.startswith(("http://", "https://")):
problems.append(f"relative URL for {code}: {href}")
if not valid_hreflang(code):
problems.append(f"invalid code: {code}")
alternates[code] = urljoin(response.url, href)
return response.url, alternates, problems
def normalise(url):
"""Compare URLs without default ports, host case or a trailing slash getting in the way."""
parts = urlsplit(url)
port = f":{parts.port}" if parts.port and parts.port not in (80, 443) else ""
return urlunsplit((parts.scheme.lower(), (parts.hostname or "") + port, parts.path.rstrip("/") or "/", parts.query, ""))
def check_return_links(url, alternates, session=None, timeout=20):
"""Fetch each alternate and report why it breaks the return-link rule. An empty result means all link back."""
session = session or requests.Session()
target = normalise(url)
problems = {}
for code, alternate in alternates.items():
if normalise(alternate) == target:
continue
final, back, _ = hreflang_map(alternate, session, timeout)
if normalise(final) != normalise(alternate):
problems[code] = f"redirects to {final}"
elif not back:
problems[code] = "has no hreflang annotations"
elif target not in {normalise(u) for u in back.values()}:
problems[code] = "does not link back"
return problems
Run against our own homepage:
import requests
from hreflang import hreflang_map, check_return_links
session = requests.Session()
session.headers["User-Agent"] = "ExampleCollector/1.0 (+https://example.com/bot)"
url, alternates, problems = hreflang_map("https://shifter.io/", session)
for code, alternate in sorted(alternates.items()):
print(code, alternate)
print("problems:", problems)
print("return-link problems:", check_return_links(url, alternates, session))
de https://shifter.io/de
en https://shifter.io/
es https://shifter.io/es
fr https://shifter.io/fr
ja https://shifter.io/ja
ko https://shifter.io/ko
pt-BR https://shifter.io/pt
x-default https://shifter.io/
zh-Hans https://shifter.io/cn
problems: []
return-link problems: {}
Two limits are worth knowing. The code reads annotations in the HTML only; sites that publish them in an HTTP header or a sitemap need those read too, and sitemaps as a discovery source covers the sitemap side. And check_return_links fetches every alternate it is given, so on a page with 150 versions, sample a few or pace the requests.
Putting it to work
- Start from one page per template. A product, a category and an article page usually share an annotation pattern, so one fetch of each shows how the whole site maps its markets.
- Key versions by the set’s own URLs. Store the URL each alternate declares, and treat redirects as findings rather than noise.
- Fetch each version as a local visitor. Use an exit in the market and that market’s language, then check you were not redirected. Location decided two of our four redirect cases, as it does in which countries get geo-blocked most.
- Validate codes before trusting them. A reversed or invented code can quietly assign a page to the wrong market.
- Normalise what you collect. Once you have every version, prices, dates and numbers still differ in format, as covered in normalising prices, numbers and dates across locales, and the comparison itself is the subject of detecting geo-personalised pricing.
The bottom line
hreflang is the closest thing the web has to a published index of a page’s markets. A third of top sites provide it, one fetch reveals the whole set, and most of the set links together as it should.
It is a map, not a guarantee. Codes can be wrong, the URL you land on may not be the one the set uses, and some versions open only for visitors in the right place or with the right language. Read the annotations, validate them, and fetch each version the way a local visitor would.
Sources and references
- Google Search Central, Tell Google about localized versions of your page.
- Tranco, list Q2K34, aggregating domain rankings from 1 to 30 September 2026.
- Babel, whose CLDR data supplies the language, script and region lists in the code.
- Homepages and alternates fetched by Shifter on 1 October 2026, directly and through Shifter exits, using the code above.