For twenty years, extracting data from web pages meant writing selectors: find the element, grab the text, fix it when the site changes. Large language models offer a different deal. Give the model the page, describe the fields you want, and get structured data back, with no selectors to write and none to break on a redesign.
The deal is real, but it has a price, a speed and a failure mode, and all three are worth measuring before rebuilding a pipeline around it. So we measured them. This piece reports a small, honest test on live pages, what the misses revealed, what it costs at scale, and where a model belongs in an extraction pipeline.
Key takeaways
- On 18 live news articles, a large Claude model extracted every headline exactly, 17 of 18 dates and 14 of 18 author lists, scored against each page’s own structured data.
- Most “misses” were not model errors. On every page where the authors disagreed, the structured data held a generic placeholder; on two of them the model reported the byline the visible page actually showed.
- The model invented nothing: every headline and author it returned appears in the page text. It still made one real mistake, naming a photographer as the author.
- It cost about 1.9 cents a page at list price and took a median of 2.4 seconds. Parsing structured data costs almost nothing and takes milliseconds.
- Use models where selectors are expensive to maintain, such as many different site templates or fields with no structured source, and validate everything they return.
What we tested
| Setting | Value |
|---|---|
| Pages | 18 live articles from the section fronts of one major news publisher, collected on 28 September 2026 |
| Model input | The page’s visible text, including navigation and other page furniture, averaging about 8,700 characters |
| Fields | Headline, publication date as shown on the page, and author names |
| Model | Claude Opus 5, at low effort, with a JSON schema constraining the output |
| Answer key | The same page’s own JSON-LD structured data |
| Scoring | Headline and date must match exactly; author lists must match as sets |
The answer key is deliberately the kind of data covered in stop parsing HTML: the structured data publishers embed for search engines. It is usually right. As it turned out, not always.
Results
| Measure | Result |
|---|---|
| Headline correct | 18 of 18 |
| Date correct | 17 of 18 |
| Authors correct | 14 of 18 |
| Extracted values found in the page text | 18 of 18 headlines, 18 of 18 author lists |
| Tokens per page | about 3,380 in, 66 out |
| Latency | median 2.4 s, range 1.8 to 10.6 s |
| Cost for the run | $0.33 at list price, about 1.9 cents per page |
The misses were the interesting part
The headline figures understate the model and overstate the answer key. Looking at each disagreement:
- An opinion column. The structured data listed the author as a generic staff reporter placeholder. The visible byline named the actual columnist. The model returned the columnist.
- A wire story. Structured data again gave the generic placeholder; the visible byline read “Staff and agencies”. The model returned what the page showed.
- A newsletter sign-up page that had slipped into the set. It showed no author and no date. The structured data supplied a generic author and a date anyway; the model left both fields empty rather than invent them.
- A photo essay. Structured data gave the same generic placeholder. The page credited the photographs to a named photographer, and the model returned the photographer as the author. That is a genuine mistake.
So of the four pages with a disagreement, two reflected the answer key being less accurate than the visible page, one was the model correctly declining to guess, and one was a real error. That changes the question worth asking. Structured data is a strong default, but it is not ground truth, and a model reading the visible page can catch where it is wrong, just as it can occasionally misread what it sees.
What it costs at scale
The test cost $0.33 for 18 pages at Claude Opus 5’s list price of $5 per million input tokens and $25 per million output tokens. That is about 1.9 cents a page, which comes to roughly $18,600 per million pages. Anthropic’s Batch API halves those token prices for work that can wait. Smaller models are cheaper per token again, Claude Haiku 4.5 is listed at $1 and $5, but we did not test them here, and their accuracy on your pages is something to measure, not assume.
Latency matters too. A median of 2.4 seconds per page, with occasional outliers, is fine for monitoring and enrichment and a real constraint for high-volume crawling. Parsing JSON-LD or running selectors takes milliseconds and costs effectively nothing beyond the fetch. Those extraction costs sit on top of collection costs, which is why cost per clean record should include both.
When to use which
| Situation | Best approach |
|---|---|
| The page publishes the fields in JSON-LD or embedded JSON | Parse the structured data |
| High volume from one or a few site templates | Selectors or structured data, with validation |
| Many different sites, each with its own template | A model, validated, possibly generating selectors you then reuse |
| Fields that exist only in prose, such as terms, eligibility or specifications | A model |
| Structured data fails validation on a page | A model as the fallback |
| Low volume, high value, frequent layout change | A model |
The strongest pipelines combine them. Parse structured data first, fall back to a model when a required field is missing or fails validation, and record which method produced each record, the same fallback chain pattern described for structured data. Where one site template covers thousands of pages, it can also pay to have a model propose selectors once and run those selectors cheaply on every page, with the model on standby for when they break.
Using a model safely
Four practices turn model extraction from impressive to dependable.
Constrain the output. A JSON schema makes the model return the fields and types you asked for; still check the stop reason, since a refused or truncated response has nothing to parse. Tell the model to leave a field empty when the page does not show it; in our test it did exactly that.
Check grounding. Every extracted string that should appear on the page, such as a name, a headline or a product title, should be found in the page text. It is a cheap check that catches invented values. It will not catch a real name attached to the wrong field, as the photographer case shows, so it is a filter, not a proof.
Sample for human review. Score a small random sample each week against a person’s reading of the page, and track accuracy per field and per site over time.
Treat page text as untrusted input. A page is written by someone else. Use models on web content for extraction only, never let page content trigger actions, and keep the model’s instructions separate from the page.
This is the extraction call we used, followed by the grounding check:
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const SCHEMA = {
type: "object",
properties: {
headline: { type: "string" },
date_published: { type: "string", description: "YYYY-MM-DD, as shown on the page" },
authors: { type: "array", items: { type: "string" } },
},
required: ["headline", "date_published", "authors"],
additionalProperties: false,
};
export async function extractArticle(pageText) {
const response = await client.beta.messages.create({
model: "claude-opus-5",
max_tokens: 2000,
betas: ["server-side-fallback-2026-07-01"],
fallbacks: "default",
output_config: { effort: "low", format: { type: "json_schema", schema: SCHEMA } },
messages: [{
role: "user",
content:
"Below is the visible text of a news article web page, including navigation and other page furniture. " +
"Extract the article's headline exactly as written, its publication date as shown on the page (YYYY-MM-DD), " +
"and the author names. Leave a field empty if the page does not show it.\n\n<page>\n" + pageText + "\n</page>",
}],
});
if (response.stop_reason === "refusal") return null;
const text = response.content.filter((b) => b.type === "text").map((b) => b.text).join("");
return { record: JSON.parse(text), usage: response.usage };
}
// Grounding check: every extracted string must actually appear in the page text.
export function grounded(record, pageText) {
const norm = (s) => s.replace(/[‘’]/g, "'").replace(/[“”]/g, '"').replace(/\s+/g, " ").toLowerCase();
const page = norm(pageText);
return {
headline: page.includes(norm(record.headline)),
authors: record.authors.every((a) => page.includes(norm(a))),
};
}
The input matters as much as the model. A model can only extract what the page actually contains, so a block page, a consent wall or an empty shell produces a confident extraction of the wrong content. Validate what you fetched before you extract from it, as covered in the silent failure rate, and fetch from the market whose version of the page you need.
The limits of this test
Eighteen pages from one publisher is a small sample, chosen to be honest rather than definitive. News articles are also an easy case: clear headlines, visible bylines, well-labelled dates. Product pages with variants, prices and availability, or pages in other languages, will behave differently. We tested one model at one effort level. The method is simple to repeat on your own pages, and your own pages are the only benchmark that matters.
The bottom line
A model reading the page got every headline right, invented nothing, and in several cases reported the page more faithfully than the page’s own structured data. It also misattributed an author once, cost cents rather than fractions of a cent, and took seconds rather than milliseconds.
That makes it a strong tool for the parts of extraction that selectors handle badly: many templates, prose-only fields and fallbacks when structured data fails. It is not a replacement for structured data where structured data exists. Parse what the page publishes, reach for a model where it does not, validate both, and record which one you used.
Sources and references
- Anthropic, Pricing. Model and Batch API prices at the time of testing.
- Schema.org, NewsArticle.
- Test run by Shifter on 28 September 2026 against 18 publicly accessible news articles, using the code above.