Most crawlers discover pages the expensive way: fetch a page, extract its links, queue them, repeat. It works, but it spends most of its budget on navigation, listings and pages you have already seen, and it can still miss pages that nothing links to.
Many sites publish a list of their URLs on purpose. XML sitemaps exist so search engines can find pages efficiently, and the same file tells a data team what exists on a site, how it is organised and, sometimes, what changed recently. This guide covers how to find and read sitemaps properly, and why the most useful field in them, lastmod, needs checking before you trust it. We checked it on two real sites, and it failed in two different ways.
Key takeaways
- A sitemap lists up to 50,000 URLs per file, and a sitemap index can list up to 50,000 sitemaps, so even very large sites can publish a complete inventory.
- Sitemaps are usually declared in robots.txt with a
Sitemap:line; check there before guessing a location. lastmodis the field that could save you the most fetches, and it is the least reliable. Google says it useslastmodonly if it is “consistently and verifiably” accurate.- In our check, one major news publisher’s news sitemap gave every entry the same timestamp, the moment the file was generated. A large document archive published over 10,000 URLs with no
lastmodat all. - Use sitemaps for discovery, verify
lastmodper site before scheduling on it, and keep link crawling as a safety net.
What the protocol allows
The sitemaps protocol is short and worth knowing precisely:
| Rule | Detail |
|---|---|
| Size limits | ”no more than 50,000 URLs” and “no larger than 50MB (52,428,800 bytes)” per sitemap file, uncompressed |
| Sitemap indexes | A file listing other sitemaps, with the same 50,000 entry and 50MB limits |
| Compression | Files may be gzip-compressed, as long as they are within the limit once uncompressed |
lastmod | Optional, in W3C Datetime format, which may be just a date, such as YYYY-MM-DD |
| Discovery | A Sitemap: line in robots.txt; a site can list several |
Two optional fields you will see, changefreq and priority, can be ignored. Google says plainly that it “ignores <priority> and <changefreq> values”, and there is little reason for anyone else to trust them either.
News sites often publish a separate news sitemap covering only recent articles, with extra fields such as a publication date. These are small, frequently regenerated and very useful for monitoring recent content.
Reading sitemaps properly
A robust reader needs to do four things: find the sitemaps from robots.txt, handle gzip, follow sitemap indexes recursively, and parse lastmod into something comparable. It should also cap how many files it fetches, since an index can point to thousands of them.
import gzip
import io
import re
import urllib.request
from datetime import datetime, timezone
from defusedxml.ElementTree import fromstring # pip install defusedxml
NS = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
UA = "Mozilla/5.0 (compatible; sitemap-reader)"
MAX_BYTES = 52_428_800 # the protocol's 50MB limit, uncompressed
def fetch(url, opener=None):
opener = opener or urllib.request.build_opener()
req = urllib.request.Request(url, headers={"User-Agent": UA})
with opener.open(req, timeout=45) as r:
body = r.read(MAX_BYTES + 1)
if body[:2] == b"\x1f\x8b":
body = gzip.GzipFile(fileobj=io.BytesIO(body)).read(MAX_BYTES + 1)
if len(body) > MAX_BYTES:
raise ValueError(f"{url} exceeds the 50MB sitemap limit")
return body
def sitemap_roots(origin, opener=None):
"""Sitemaps declared in robots.txt, falling back to /sitemap.xml."""
try:
robots = fetch(origin.rstrip("/") + "/robots.txt", opener).decode("utf-8", "replace")
found = re.findall(r"(?im)^\s*sitemap:\s*(\S+)", robots)
except OSError:
found = []
return found or [origin.rstrip("/") + "/sitemap.xml"]
def parse_lastmod(value):
if not value:
return None
value = value.strip().replace("Z", "+00:00")
try:
dt = datetime.fromisoformat(value)
except ValueError:
return None
return dt if dt.tzinfo else dt.replace(tzinfo=timezone.utc)
def walk(url, opener=None, max_files=20, _seen=None):
"""Yield (loc, lastmod) from a sitemap or sitemap index, following nested indexes."""
seen = _seen if _seen is not None else set()
if url in seen or len(seen) >= max_files:
return
seen.add(url)
root = fromstring(fetch(url, opener))
if root.tag == NS + "sitemapindex":
children = sorted(root.findall(NS + "sitemap"),
key=lambda s: parse_lastmod(s.findtext(NS + "lastmod")) or datetime.min.replace(tzinfo=timezone.utc),
reverse=True)
for child in children:
yield from walk(child.findtext(NS + "loc").strip(), opener, max_files, seen)
else:
for u in root.findall(NS + "url"):
yield u.findtext(NS + "loc").strip(), parse_lastmod(u.findtext(NS + "lastmod"))
def changed_since(origin, since, opener=None, max_files=20):
"""URLs whose sitemap lastmod is newer than `since`, plus those with no lastmod at all."""
changed, undated, total = [], [], 0
for root in sitemap_roots(origin, opener):
for loc, lastmod in walk(root, opener, max_files):
total += 1
if lastmod is None:
undated.append(loc)
elif lastmod > since:
changed.append((loc, lastmod))
return {"total": total, "changed": changed, "undated": undated}
The index walker visits child sitemaps newest first where the index provides dates, so a capped run reads the most recently updated parts of a large site first. Two details are there for safety. A sitemap is untrusted input from someone else’s server, so the code parses it with defusedxml, which refuses the entity tricks that can make standard XML parsers expand a small file into gigabytes, and it caps both the download and the decompressed size at the protocol’s 50MB limit, so a compressed file cannot inflate without bound. The opener parameter lets you route requests through a proxy when you need to see a market-specific version of a site, the same approach as for any other fetch.
Trust lastmod only after checking it
lastmod promises exactly what incremental collection needs: a list of what changed since the last run. If it were reliable everywhere, a crawler could fetch only the changed pages and skip everything else. It is not reliable everywhere, and our two test sites show both common failures.
Failure one: lastmod is the generation time. We read the news sitemap of a major news publisher. Every dated entry carried one of two timestamps, a second apart: the moment the file was generated, not the moment each article changed. Filtering for “changed in the last six hours” returned every URL in the file. Used naively, that lastmod would trigger a refetch of everything, every time.
Failure two: no lastmod at all. We read the sitemap of a large technical document archive. It listed 10,236 URLs, and none carried a lastmod. That is a perfectly good discovery source and no help at all for deciding what to refetch.
This is why Google’s own guidance says it uses lastmod only “if it’s consistently and verifiably (for example by comparing to the last modification of the page) accurate”, and why you should apply the same test yourself. Before scheduling on a site’s lastmod:
- Check the spread. If most entries share one timestamp, or the values track the file’s generation time, the field is not describing page changes.
- Sample and compare. Fetch a sample of pages and compare
lastmodwith an independent signal, such as the page’s owndateModifiedin structured data or a content fingerprint. How to extractdateModifiedis covered in stop parsing HTML, and how to compare content without drowning in noise in change detection at scale. - Score the site. Record, per site, how often
lastmodchanged when the content did, and how often the content changed whenlastmoddid not. Uselastmodfor scheduling only on sites that pass, and keep checking, because a site’s sitemap generator can change without notice.
Where sitemaps fit in a collection pipeline
Sitemaps are strongest as the first step, not the only one:
- Discovery. A sitemap gives you the inventory directly, including pages that are poorly linked. For catalogue and content sites, it is often the most complete URL list available.
- Segmentation. Sitemaps are commonly split by content type or section, such as products, categories, articles and videos. The split tells you how a site is organised before you fetch a single page.
- Incremental collection, on sites whose
lastmodpasses the checks above, and via news sitemaps for recent articles. - Coverage checks. Comparing what your crawler found with what the sitemap lists shows what you are missing.
Keep link-based crawling as a safety net. Sitemaps can be stale, incomplete or deliberately limited to what a site wants indexed. And respect the site’s robots.txt for the pages themselves: a URL appearing in a sitemap is an invitation to search engines to index it, not a waiver of anything else the site has asked, as covered in robots.txt, AI opt-outs and reservation signals.
Used this way, sitemaps cut the fetch budget in the two places it is usually wasted: navigating to find pages, and refetching pages that have not changed. The second saving depends entirely on lastmod being trustworthy, which is why cost-aware crawl scheduling should treat an unverified lastmod as a hint, not a fact.
The bottom line
Sitemaps are the cheapest URL inventory on the web: a site telling you, in a standard format, what it wants found. Read them from robots.txt, handle gzip and nested indexes, cap what you fetch, and use them first for discovery.
Then treat lastmod the way Google does: useful only when it has been shown to be accurate for that site. On one site we checked, it was just the time the file was generated. On another, it did not exist. On a site where it holds up, it can cut refetching dramatically. The only way to know which kind of site you are dealing with is to check.
Sources and references
- Sitemaps.org, Sitemaps XML format.
- Google Search Central, Build and submit a sitemap.
- Sitemaps read by Shifter on 29 September 2026 using the code above.