Knowledge

Public-Record Collection for Investigative Reporting: Method, Ethics and Verification

Public records drive investigations, but only if collected carefully, preserved as served and checked before publication. A method for newsrooms.

James Meadow

October 5, 2026 · 8 min read

Many of the most consequential investigations start in public records: a company registry that shows the same director behind fifty shell companies, procurement notices that reveal a contract was split to stay under a threshold, court filings that contradict a public statement. The information was public all along. The work was in collecting it systematically, keeping it intact, and checking it before anyone relied on it.

Doing that at scale is a data collection problem, and the habits that make collection reliable for businesses matter even more when the result is published under a byline. This guide sets out a method for collecting public records for reporting and research: where to start, how to collect, how to preserve evidence, how to verify it, and the ethical questions that come with it.

Key takeaways

  • Start with official bulk data and APIs. Many registries publish free downloads or APIs that are more complete and more defensible than scraping their websites.
  • Collect with a written plan: what you need, from where, how often, and what you will not collect.
  • Preserve every record exactly as served, with a hash, a timestamp and the address it came from, so you can prove later what the source said.
  • Verify before you publish: confirm records against a second source or the issuing body, and treat any single record as a lead, not a finding.
  • Public does not mean harmless. Personal data in public records still deserves minimisation, care and a clear public-interest reason.

Start with what is published for reuse

Before writing a scraper, check whether the source already publishes its data for reuse. It often does, and official data is easier to defend than scraped copies:

SourceWhat is availableNotes
UK Companies HouseA free public data API and free bulk downloads, including a monthly company snapshot, accounts data and a file of people with significant controlOfficers, filings and ownership for UK companies
EU procurement, TEDNotices from the Official Journal supplement, around 800,000 a year, available as open data and through an APITenders and contract awards across the EU
US federal courts, PACERElectronic court records at $0.10 a page, capped at $3 per document, with fees waived for users who spend $30 or less in a quarterDockets and filings; costs add up at scale

Bulk data has another advantage: it captures the whole register at a point in time, which lets you find patterns, such as clusters of companies at one address, that page-by-page browsing never shows. Where only a website exists, collect from it carefully, as below.

Collect with a plan

A short written plan before collection starts keeps the work focused and defensible:

  • The question. What you are trying to establish, and which records could confirm or refute it.
  • The sources. Which registers, portals and archives, and why each is authoritative for its part.
  • The fields. What you will collect, and what you will deliberately leave out, especially personal data you do not need.
  • The cadence. One snapshot, or repeated collection to catch changes and deletions.
  • The method. Official API, bulk download or careful collection from web pages, at a pace that places no meaningful load on public services.

Repeated collection is often where stories are found. Records that change quietly, filings that are amended, and entries that disappear are all visible only if you collected before and after. The same techniques used for change detection at scale and monitoring a supplier’s public footprint apply directly.

Preserve records as served

A screenshot is not enough evidence when a source later disputes what it published. Keep the record exactly as the server returned it, and the context needed to show where and when it came from. The function below saves the response body unchanged, named by its SHA-256 hash, and appends a log line with the URL, status, time and hash:

import hashlib
import json
from datetime import datetime, timezone
from pathlib import Path

import requests


def capture(url, folder, session=None, note=""):
    """Save a public record exactly as served, with a hash and the context needed to verify it later."""
    session = session or requests.Session()
    response = session.get(url, timeout=30)
    body = response.content
    digest = hashlib.sha256(body).hexdigest()
    retrieved = datetime.now(timezone.utc).isoformat(timespec="seconds")
    folder = Path(folder)
    folder.mkdir(parents=True, exist_ok=True)
    (folder / f"{digest}.body").write_bytes(body)
    record = {
        "requested_url": url,
        "final_url": response.url,
        "status": response.status_code,
        "retrieved_utc": retrieved,
        "sha256": digest,
        "bytes": len(body),
        "content_type": response.headers.get("content-type", ""),
        "last_modified": response.headers.get("last-modified"),
        "note": note,
    }
    with open(folder / "log.jsonl", "a", encoding="utf-8") as log:
        log.write(json.dumps(record) + "\n")
    return record

Used on a public company page from the UK register:

import requests

from capture import capture

session = requests.Session()
session.headers["User-Agent"] = "ExampleNewsroom/1.0 (+https://example.org/research)"

record = capture("https://find-and-update.company-information.service.gov.uk/company/00445790",
                 "records", session, note="registered office check")
print(record["status"], record["sha256"], record["retrieved_utc"])

Two captures of that page ten seconds apart produced the same hash, which is the point: if a record is unchanged, anyone can confirm it byte for byte, and if it changes, the hash shows exactly when. For larger collections, the WARC format described in keeping the raw response stores requests and responses together with standard tooling, and court-ready evidence from online listings covers the extra steps when material may end up in legal proceedings. A capture in a public web archive, made at the same time, adds independent corroboration.

Record where the collection ran from as well. Some portals show different content or language by visitor location, and noting the vantage point, as argued in the case for a vantage-point standard, lets others reproduce what you saw.

Verify before you publish

Public records are often wrong, out of date or ambiguous. Registers record what was filed, not what is true; names are reused; addresses are shared by many companies registered through the same formation agent. Verification turns a record into a finding:

CheckWhat it guards against
Confirm with the issuing body or a second official sourceTranscription errors and outdated entries
Match on identifiers, not namesDifferent people or companies with the same name
Check dates and versionsAmended filings, superseded records, time zone confusion
Look for innocent explanationsFormation-agent addresses, nominee directors, routine filings
Ask the subject before publicationContext the records do not show, and fairness
Keep the chain from claim to recordEvery published statement traced to a preserved, hashed record

The last row is the discipline that matters most. Every factual claim in a story should point to a specific preserved record, and anyone on the team should be able to open it and see what it said on the day it was collected.

Ethics

Collecting public records raises the same ethical questions as any reporting. The Society of Professional Journalists’ code organises them under four principles: seek truth and report it, minimize harm, act independently, and be accountable and transparent. A few apply with particular force to bulk collection:

  • Public is not the same as harmless. Registers contain home addresses, dates of birth and family details. Collect only what the story needs, restrict access to the rest, and think hard before publishing personal details of people who are not the subject.
  • Aggregation creates new information. Combining records can reveal things no single record does, which is often the point of the investigation and also the risk. Weigh the public interest of what the combination shows.
  • Collect respectfully. Public services are paid for by the public. Use bulk data where it exists, pace collection, identify your collector honestly, and do not work around access controls.
  • Mind the law where you work. Data protection, copyright and computer misuse rules apply to journalists too, and many jurisdictions provide specific journalistic or research exemptions with their own conditions. Our overviews of whether web scraping is legal and GDPR and data collection are a starting point; consult your newsroom’s counsel on specific projects.

The bottom line

Public records are one of the most powerful sources an investigation can have, and one of the easiest to misuse. Start with the official bulk data and APIs, collect to a written plan, preserve every record exactly as served with a hash and a timestamp, verify each finding against a second source and the subject, and keep personal data to what the story needs.

The result is reporting that holds up: every claim traceable to a record anyone can check, collected in a way the newsroom can explain.

Sources and references

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started