Open a modern web page, watch the network activity, and you will often see the data you want arrive as JSON before the page draws it. The prices, listings or results that a scraper would laboriously extract from HTML are sitting in a clean, structured response the page fetched for itself.
Collecting from that response instead of the rendered page can be smaller, more stable and far easier to parse. But it is not as simple as copying a URL, and two things surprise most teams: how little of a page’s JSON is actually the data, and how often the data endpoint refuses to work outside the page that called it. This guide shows how to find the right endpoint, what we measured on real pages, and how to collect from it without overstepping.
Key takeaways
- Most JSON a page loads is not data. On an airline’s flight search page, the fares came in one 44.9 KB response, about 8% of the page’s JSON and about 1.2% of its 3.6 MB total transfer.
- Some pages have no data API at all. On a national weather service’s forecast page, every JSON response belonged to the cookie-consent tool, and the forecast was in the HTML.
- Find the endpoint by searching response bodies for a value you can see on the page, not by guessing from URLs.
- Page APIs are often tied to the page’s session. The airline’s fare endpoint worked inside the page and returned an HTTP 409 when called directly.
- Use only public endpoints the page calls for anonymous visitors, at the pace a person browsing would, and prefer an official API wherever one exists.
What we measured
We loaded two public pages in a real browser on 30 September 2026 and recorded every response.
| Airline flight search | Weather forecast | |
|---|---|---|
| Requests | 61 | 46 |
| Total transferred | 3.62 MB | 2.57 MB |
| JSON responses | 15, 534 KB | 3, 1.04 MB |
| JSON responses carrying the data | 1, 44.9 KB | 0 |
| Where the data lived | A fares endpoint | The server-rendered HTML |
On the airline page, the largest JSON responses were not fares at all. A third-party feature-flag service returned the same 97 KB response four times, and a translation bundle added another 92 KB. The fares arrived in a single response from the airline’s own booking API.
On the weather page, all three JSON responses came from the consent management tool, one of them an 860 KB vendor list, while the forecast itself was rendered into the HTML on the server. “Collect from the JSON” would have collected nothing useful.
Finding the endpoint that carries the data
The reliable method is to search, not guess. Pick a value you can see on the page, such as a price, a product name or an identifier, load the page in a real browser, and find which JSON response contains it.
By hand, this is the browser’s developer tools: open the Network tab, filter to Fetch/XHR, reload, and search response bodies for the value. Automated, it is a short script:
import { chromium } from 'playwright';
// Load a page once and report which JSON responses contain a value you can see on it.
export async function findEndpoint(url, needle, waitMs = 20000) {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
const hits = [];
page.on('response', async (response) => {
const type = response.request().resourceType();
const contentType = response.headers()['content-type'] || '';
if (!['xhr', 'fetch'].includes(type) || !contentType.includes('json')) return;
try {
const body = await response.text();
if (body.includes(needle)) {
hits.push({
method: response.request().method(),
status: response.status(),
bytes: Buffer.byteLength(body),
url: response.url(),
});
}
} catch {
// Some responses (redirects, aborted requests) have no readable body.
}
});
// Analytics beacons keep many pages from ever going network-idle,
// so wait for the DOM, then until a match appears or the time runs out.
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 90000 });
for (let waited = 0; waited < waitMs && hits.length === 0; waited += 500) {
await page.waitForTimeout(500);
}
await page.waitForTimeout(1000);
await browser.close();
return hits;
}
Given the airline’s flight page and a fare visible on it, the script returned exactly one match out of 61 requests: the booking API’s availability response. Note the wait strategy. Our first version waited for the page to go network-idle, which it never did, because analytics beacons kept firing; it timed out after 90 seconds. Waiting for the DOM and then for a match is more dependable.
Can you call it directly?
Often not. When we requested the same airline availability endpoint directly, without the page, it returned HTTP 409 for requests from six different countries alike, while the page itself loaded it without trouble. Page APIs commonly depend on things the page sets up first: session cookies, tokens embedded in the HTML, request headers, or a sequence of earlier calls.
That leaves two approaches:
- Let the page make the call. Load the page in a browser and read the JSON response as it arrives, as the script above does. You get the clean data without reconstructing the session, at the cost of running a browser.
- Replay only simple, public endpoints. Some endpoints need nothing but a URL. The same airline also publishes a public fares endpoint that answers plain requests, which we used in a separate test of whether prices change by exit country. Where an endpoint works that simply, it is the cheapest option by far.
What not to do is work around the session: harvesting tokens, forging headers or replaying authenticated calls to reach data the site did not make available to anonymous visitors. That is where collection stops being observation.
Why JSON is worth it when it is there
When the data does arrive as JSON, the gains are real:
- Size. The fares response was 44.9 KB against 3.6 MB for the full page. On bandwidth-billed collection, that difference is most of the bill, as cost per clean record shows.
- Structure. Fields arrive named and typed, with no selectors to write or maintain.
- Completeness. Responses often carry more than the page displays, such as identifiers, all fare classes or stock flags.
The trade-offs are just as real. Undocumented endpoints change without notice, and a response that changes shape breaks your parser silently, which is exactly the problem schema drift monitoring exists for. A version number in the path, like the v4 in the airline’s booking API, is a mild signal of stability, not a promise.
Where this fits among the alternatives
| Source | Stability | Effort | Use when |
|---|---|---|---|
| Official, documented API | Highest | Lowest | Always check first |
| JSON-LD or embedded page state | High | Low | The page publishes it, see stop parsing HTML |
| The page’s own JSON endpoints | Medium | Medium | Data loads after the page, and the endpoint is public |
| Rendered HTML and selectors | Lowest | Highest | Nothing else carries the data |
| A model reading the page | Varies | Low per site, high per page | Many templates, as in LLM extraction vs selectors |
Collect responsibly
The endpoints a page calls are part of the site, and the site’s rules still apply. Use only endpoints the page calls for anonymous visitors, pace requests the way a person browsing would, respect robots.txt and the site’s terms, and leave anything behind a login or a token alone unless you have permission. Where a site offers an official API, use it; it is more stable, and it is what the site has agreed to support. The broader principles are in robots.txt, AI opt-outs and reservation signals.
The bottom line
The data behind a modern page often arrives as JSON, and when it does, collecting it there is smaller, cleaner and easier to maintain than parsing HTML. But finding it takes a search, not a guess: on the pages we measured, the useful JSON was one response among many, and on one page it did not exist at all.
Search response bodies for a value you can see, check whether the endpoint works on its own, let the page make the call when it does not, and stay within what the site offers to any anonymous visitor.
Sources and references
- Pages loaded and measured by Shifter on 30 September 2026, using the code above.
- Playwright, Network events documentation.