Most scrapers are built the same way: open the page, find the element that holds the price, write a CSS selector, repeat for every field. It works until the site ships a redesign, renames a class, or wraps the price in a new component. Then the selector returns nothing, or worse, returns the wrong thing, and the pipeline keeps running.
Many pages already publish the same data in a form meant for machines. Search engines asked for it, sites supplied it, and it has been sitting in the page source all along. This tutorial shows how to find it, extract it with a few dozen lines of code, and handle the cases where it is missing or incomplete, which, as a real example below shows, happens more often than the documentation suggests.
Key takeaways
- JSON-LD appeared on 41% of pages in the 2024 Web Almanac, up from 34% in 2022. For products, articles, events and organisations, it is often the most stable source on the page.
- Structured data changes far less often than page layout, because sites depend on it for search results. Selectors break on redesigns; structured data usually survives them.
- It is not always complete. A product page can publish full records for the variant on screen and only bare links for every other variant. Validate every record.
- The most robust extractor tries structured data first, embedded JSON second, and CSS selectors last, and records which one it used.
What “structured data” means on a web page
There are four places machine-readable data commonly lives:
| Source | What it looks like | Typical contents |
|---|---|---|
| JSON-LD | <script type="application/ld+json"> blocks | schema.org objects: Product, Offer, NewsArticle, Organization, Event, BreadcrumbList |
| Microdata | itemprop attributes on visible elements | The same schema.org vocabulary, spread across the markup |
| Open Graph and meta tags | <meta property="og:..."> | Title, description, image, sometimes price |
| Embedded application state | A large JSON object the page’s JavaScript reads, such as __NEXT_DATA__ | Often everything the page displays, and more |
The HTTP Archive’s 2024 Web Almanac measured their spread across the web: JSON-LD grew “from 34% in 2022 to 41% in 2024,” microdata held steady at 26%, and RDFa and Open Graph, which include the social sharing tags most sites add, appeared on 66% and 64% of pages. JSON-LD is the one to reach for first, because it is a self-contained block of data rather than attributes scattered through the layout.
Why it is more reliable than selectors
A CSS selector depends on how a page looks. Structured data depends on what a page means. Sites change how their pages look constantly. They change what their structured data says far less often, because it feeds rich results in search engines, and breaking it has a visible cost to the site.
It also fails more honestly. A selector that stops matching can silently match a different element and return a plausible wrong value, the kind of error described in the silent failure rate. A JSON-LD object either contains a price field or it does not, which makes validation straightforward.
Extracting JSON-LD in Python
The standard library is enough. This extractor collects every JSON-LD block, tolerates broken ones, and walks the containers that sites use to nest objects: lists, @graph, and hasVariant, which schema.org uses for product variants.
import json
from html.parser import HTMLParser
class _Collector(HTMLParser):
"""Collect JSON-LD blocks and embedded JSON state from an HTML page."""
def __init__(self):
super().__init__()
self.blocks, self._buf, self._kind = [], None, None
def handle_starttag(self, tag, attrs):
a = dict(attrs)
if tag == "script" and a.get("type", "").lower() == "application/ld+json":
self._buf, self._kind = [], "json-ld"
elif tag == "script" and a.get("id") == "__NEXT_DATA__":
self._buf, self._kind = [], "next-data"
def handle_data(self, data):
if self._buf is not None:
self._buf.append(data)
def handle_endtag(self, tag):
if tag == "script" and self._buf is not None:
raw = "".join(self._buf).strip()
try:
self.blocks.append((self._kind, json.loads(raw)))
except json.JSONDecodeError:
self.blocks.append((self._kind + "-invalid", raw[:200]))
self._buf = self._kind = None
def _walk(node):
"""Yield every JSON-LD object, flattening lists and nested containers."""
if isinstance(node, list):
for item in node:
yield from _walk(item)
elif isinstance(node, dict):
yield node
for key in ("@graph", "mainEntity", "itemListElement", "hasVariant"):
if key in node:
yield from _walk(node[key])
def _types(obj):
t = obj.get("@type", [])
return {t} if isinstance(t, str) else set(t)
def jsonld_objects(html, wanted_type=None):
"""Every JSON-LD object on the page, optionally filtered by schema.org type."""
collector = _Collector()
collector.feed(html)
objs = [o for kind, data in collector.blocks if kind == "json-ld" for o in _walk(data)]
return [o for o in objs if wanted_type is None or wanted_type in _types(o)]
On a live Guardian article, jsonld_objects(html, "NewsArticle") returns the headline, the publication and modification timestamps, and the author, with no selectors at all. Those timestamps alone are worth the effort: they are exact, machine-readable and consistent across every article on the site.
Normalising products
Products need a little more care, because prices live inside nested offers, sometimes as a single Offer and sometimes as an AggregateOffer with a price range.
def products(html):
"""Return normalised product records found in a page's JSON-LD."""
out = []
for obj in jsonld_objects(html, "Product"):
offers = obj.get("offers") or {}
offer = offers[0] if isinstance(offers, list) and offers else offers
if isinstance(offer, dict) and "AggregateOffer" in _types(offer):
price = offer.get("lowPrice")
else:
price = offer.get("price") if isinstance(offer, dict) else None
brand = obj.get("brand")
out.append({
"name": obj.get("name"),
"sku": obj.get("sku") or obj.get("gtin13") or obj.get("mpn"),
"brand": brand.get("name") if isinstance(brand, dict) else brand,
"price": float(price) if price not in (None, "") else None,
"currency": offer.get("priceCurrency") if isinstance(offer, dict) else None,
"availability": (offer.get("availability") or "").rsplit("/", 1)[-1] if isinstance(offer, dict) else None,
})
return out
Run against a page with a standard product and offer, it returns clean records such as {"name": "Trail Runner", "sku": "TR-01", "brand": "Acme", "price": 89.0, "currency": "EUR", "availability": "InStock"}, regardless of how the page is styled.
When structured data is incomplete: a real example
Documentation examples make this look easy. Real pages are messier, and it is worth showing one.
We ran the extractor against the product page for a popular shoe on a large Shopify store. The page’s JSON-LD described a ProductGroup, the schema.org type for a product sold in variants, with the product name, brand, description and images, and 49 variants. Only 7 of them, the sizes of the colour on screen, were complete products with a price of $100.00 and an availability status. The other 42, every other colour and size, were bare references: a type and a URL, nothing else. The review rating sat in a second, separate JSON-LD block.
So the structured data described the page, not the catalogue. A pipeline that assumed “all variants are in the JSON-LD” would have silently priced one colour and recorded nothing for the rest.
The same store also exposes a public JSON representation of each product, which matched: seven variants for that colour, each with a price and an availability flag. But there the price came as 10000, in minor units. A pipeline that mixed the two sources without normalising would have recorded one shoe at $100 and the same shoe at ten thousand dollars.
Three lessons follow, and they apply well beyond Shopify:
- Validate, do not assume. A JSON-LD block that parses is not a complete record. Check that every field you need is present, and count what you got against what you expected.
- Follow the references when structured data is partial. Variant URLs, embedded JSON and public product endpoints often fill the gaps, and they are usually cleaner than the page.
- Normalise units explicitly. Minor units, price ranges, tax-inclusive and tax-exclusive prices, and currency all need handling in code, not by assumption.
A fallback chain that records its source
Put the pieces together as a chain: try structured data, then embedded JSON, then selectors, and keep a note of which one produced each record.
REQUIRED = ("name", "price", "currency")
def extract_product(html, embedded=None, css_fallback=None):
for source, candidates in (
("json-ld", products(html)),
("embedded-json", embedded(html) if embedded else []),
("css", css_fallback(html) if css_fallback else []),
):
for record in candidates:
if all(record.get(f) not in (None, "") for f in REQUIRED):
return {**record, "source": source}
return None
The source field earns its place quickly. When a site that always yielded json-ld records starts yielding css ones, its structured data has changed or disappeared, and you want to know that before the selector fallback breaks too. It is also a useful input to a target health score.
Doing it with the Web Scraping API
If you fetch pages through Shifter’s Web Scraping API, the same approach works without running a browser yourself. The API’s extract_rules parameter maps CSS selectors to JSON fields, and its html output returns an element’s inner HTML, so a rule that selects the JSON-LD script returns the raw block for you to parse:
{
"jsonld": { "selector": "script[type='application/ld+json']", "output": "html" }
}
A single rule returns the first matching element, so for pages with several JSON-LD blocks, request the full HTML and run the extractor above on it. Combine either approach with render_js=1 for pages that inject their structured data with JavaScript, and use auto_parser=1 when you are fetching a JSON endpoint directly, such as a product JSON URL, to get the parsed body back. Missing fields come back as null rather than failing the request, which fits the validate-then-fall-back pattern above. The full syntax is in the extraction rules documentation.
What structured data will not give you
Structured data describes what the site chose to publish for search engines. It may lag the visible page, omit fields the site does not care to expose, or describe the default variant rather than the one on screen. For prices in particular, compare it against the visible page on a sample basis, because a stale JSON-LD price and a current on-page price are both “correct” from different points of view. And for sites that vary content by visitor location, structured data varies too, so capture it from the market you care about, as covered in scraping flight and hotel prices.
The bottom line
Before writing another selector, open the page source and search for application/ld+json. On a large share of the web, the data you want is already there, labelled with a shared vocabulary, and far less likely to change than the layout around it.
Extract it first, validate it, fall back to embedded JSON and then to selectors when it is incomplete, and record which source every record came from. Your extractors will break less often, and when they do, they will tell you.
Sources and references
- HTTP Archive, Web Almanac 2024: Structured Data, 11 November 2024.
- Schema.org, Product, ProductGroup, Offer and AggregateOffer.
- Shifter, Web Scraping API extraction rules and rendering JavaScript documentation.