Knowledge

Web Data Collection for Academic Research: Reproducibility, Ethics Review and How to Cite a Scraped Dataset

Scraped data is now common in research, but often hard to reproduce, review or cite. How to collect web data that survives peer review, ethics boards and reuse.

James Meadow

October 6, 2026 · 10 min read

Web data has become ordinary research material. Social scientists study platform discourse, economists track online prices, computer scientists build corpora, and public health researchers follow misinformation. Much of it is collected with scrapers, and much of it shares the same weaknesses: nobody outside the team can reproduce the collection, the ethics review was written for surveys, and the dataset is cited as “data collected by the authors” or not at all.

None of that is inevitable. A few practices, most of them cheap, make scraped data reproducible, reviewable and citable. This guide covers how to design a collection that survives peer review, what an ethics board needs to see, how to package and cite the dataset, and a small tested script that produces the citation and manifest files.

Key takeaways

  • The web changes and differs by viewer, so a scraped dataset cannot be reproduced by re-running the scraper. Reproducibility means recording exactly what was collected, when, from where and how, and preserving it.
  • Write the collection protocol before collecting, and treat changes to it as amendments, as you would for any other method.
  • Ethics review for web data turns on expectations, harm and identifiability, not just on whether the data is public. Bring a data management plan, not just a consent waiver.
  • Package the dataset with a manifest of file hashes and collection context, a licence, and a CITATION.cff file, then deposit it where it gets a persistent identifier.
  • Cite the dataset as a research output in its own right, with a version and an identifier, alongside the paper.

Why scraped data is hard to reproduce

Re-running the same code a year later does not produce the same data. Pages change, disappear and move. Many sites show different content by country, language, device or login state. Platforms change their markup and their access rules. And collection itself is lossy: blocks, timeouts and rate limits drop records unevenly, often in ways that correlate with the very things being studied.

So reproducibility for web data rests on three things rather than on re-collection:

  1. Documentation of the method in enough detail that someone could attempt the same collection and understand why their result differs.
  2. Preservation of what was actually collected, ideally the raw responses as well as the extracted data.
  3. Accounting for what was not collected: failures, exclusions and filtering, with counts.

Design the collection like an instrument

Treat the scraper as a measurement instrument and write its protocol before collection starts:

Protocol elementWhat to record
Population and samplingWhat the sources are, how they were chosen, and what is in and out of scope
Collection windowStart and end dates and times in UTC, and the schedule for repeated collection
Vantage pointThe country, network type, language and device profile the collector presented, since many sites vary by these
ToolingCode repository and commit, library versions, browser version if one was used
Access rulesHow robots.txt, terms of service and rate limits were handled, and the request pace
Failure handlingRetries, what counts as a failed record, and how failures are logged
ProcessingExtraction, cleaning, deduplication and filtering steps, each with counts before and after

Two of these are routinely missing from papers. The vantage point matters because the same URL can return different content to different visitors, a problem we argue should be recorded as standard in the case for a vantage-point standard. And failure counts matter because a dataset missing 8% of its records is only interpretable if readers know which 8%.

Where the study allows, keep the raw responses as well as the extracted data. Storing them in the standard WARC format, as described in keeping the raw response, lets you and later researchers re-extract from exactly what was collected when the parser turns out to have been wrong.

Ethics review

Many institutional ethics processes were designed around consent from participants, and web data does not fit neatly. Researchers often argue that public data needs no review; boards often respond with requirements designed for interviews. A better conversation starts from the questions that actually matter. The Association of Internet Researchers’ ethics guidelines, approved in their third version in 2019, frame internet research ethics around the context in which data was shared and the potential for harm across the whole research lifecycle, rather than around a public or private label.

Questions an ethics application for web data should answer:

  • Whose data is it, and what did they expect? A company’s price list, a politician’s public statements and a teenager’s forum post are all “public”, with very different expectations.
  • Can people be identified, directly or by combination? Quotes can be searched back to their authors; aggregating sources can identify people no single source would.
  • What harm could follow, and to whom? Consider exposure, harassment and discrimination, especially for vulnerable groups or sensitive topics.
  • How will data be minimised and protected? What is not collected, how identifiers are pseudonymised, who has access, where it is stored, and when it is deleted.
  • What will be published? Aggregates, paraphrased quotes, or raw records; and whether a shared dataset needs access controls.

Terms of service and law are separate questions from ethics, and both need an answer. Fiesler, Beard and Keegan examined the data collection provisions in the terms of service of more than a hundred social media sites, in a study presented at ICWSM in 2020, and argued that terms of service alone are a poor basis for ethical decisions in either direction. On the legal side, research uses often have specific provisions; EU copyright law, for example, includes an exception for text and data mining by research organisations, and data protection law applies to personal data whether or not it is public. Our overviews of whether web scraping is legal and GDPR and data collection are a starting point; your institution’s legal and data protection officers are the authority for your project.

Robots.txt deserves a sentence in every protocol. It is not a law, but it is the clearest machine-readable statement of what a site wants, and honouring it is easy to defend; robots.txt, AI opt-outs and reservation signals covers what the signals mean.

Package the dataset

A dataset that can be cited and reused needs more than the data files:

  • A manifest listing every file with its size and a cryptographic hash, plus the collection context from the protocol. Hashes let anyone confirm they have the exact files the paper used.
  • A README describing the fields, the method, known gaps and how to use the data responsibly.
  • A licence for what you are allowed to share, and a clear statement of what you could not share and why.
  • A CITATION.cff file, the plain-text Citation File Format that GitHub, Zenodo and reference managers read, so anyone can cite the dataset correctly with one click.

The function below writes the manifest and the CITATION.cff from a folder of data files and a description of the collection:

import hashlib
import json
from datetime import date
from pathlib import Path


def sha256(path):
    digest = hashlib.sha256()
    with open(path, "rb") as f:
        for block in iter(lambda: f.read(1 << 20), b""):
            digest.update(block)
    return digest.hexdigest()


def write_dataset_package(folder, title, authors, collection, version="1.0.0"):
    """Write a manifest (what was collected, how, and file hashes) and a CITATION.cff for a scraped dataset."""
    folder = Path(folder)
    files = sorted(p for p in folder.rglob("*") if p.is_file() and p.name not in ("MANIFEST.json", "CITATION.cff"))
    manifest = {
        "title": title,
        "version": version,
        "collection": collection,  # sources, window, vantage point, tool and code version, robots policy
        "files": [{"path": str(p.relative_to(folder)), "bytes": p.stat().st_size, "sha256": sha256(p)} for p in files],
    }
    (folder / "MANIFEST.json").write_text(json.dumps(manifest, indent=2) + "\n", encoding="utf-8")

    lines = [
        "cff-version: 1.2.0",
        'message: "If you use this dataset, please cite it as below."',
        "type: dataset",
        f"title: {json.dumps(title)}",
        f"version: {json.dumps(version)}",
        f"date-released: {date.today().isoformat()}",
        "authors:",
    ]
    for person in authors:
        lines.append(f"  - family-names: {json.dumps(person['family'])}")
        lines.append(f"    given-names: {json.dumps(person['given'])}")
        if person.get("orcid"):
            lines.append(f"    orcid: {json.dumps(person['orcid'])}")
    abstract = (f"Collected {collection['window']} from {', '.join(collection['sources'])}. "
                "See MANIFEST.json for method and file hashes.")
    lines.append(f"abstract: {json.dumps(abstract)}")  # JSON strings are valid YAML, so quotes cannot break the file
    (folder / "CITATION.cff").write_text("\n".join(lines) + "\n", encoding="utf-8")
    return manifest

Used on a small example dataset:

from package import write_dataset_package

manifest = write_dataset_package(
    "dataset",
    title="Robots.txt rules for AI crawlers on the top 10,000 domains",
    authors=[{"family": "Example", "given": "Researcher", "orcid": "https://orcid.org/0000-0002-1825-0097"}],
    collection={
        "sources": ["robots.txt files of the Tranco top 10,000 domains"],
        "window": "on 6 October 2026",
        "vantage_point": "one network in Romania",
        "tool": "survey.py at commit 3f2a9c1",
        "robots_policy": "only robots.txt itself was requested",
    },
)
print(len(manifest["files"]), "files listed")

The ORCID shown is ORCID’s own published example identifier; use yours. We validated the generated file with cffconvert, the format’s reference tool, which reported “Citation metadata are valid according to schema version 1.2.0”, and confirmed it stays valid when titles and names contain quotes and apostrophes, which naive templates break on. The tool’s APA-style output for the example reads: “Example R. (2026). Robots.txt rules for AI crawlers on the top 10,000 domains (version 1.0.0).”

Deposit and cite

A dataset in a lab folder or a personal GitHub repository can disappear. Deposit it in a repository that issues a persistent identifier:

  • General repositories such as Zenodo, operated by CERN, issue a DOI for every upload, at no cost, with a standard limit of 50 GB per record.
  • Institutional and subject repositories may be required by your funder or offer better discovery in your field.
  • Restricted access is an option when data cannot be fully public: deposit the manifest and documentation openly, and the data itself under controlled access.

The FAIR principles, published in Scientific Data in 2016, summarise what the deposit should achieve: the data should be findable, accessible, interoperable and reusable, by people and by machines.

Then cite the dataset in the paper as a reference of its own, with authors, year, title, version, repository and identifier, not only as a sentence in the methods section. Cite the exact version you analysed, and publish a new version, with a new identifier, when the data changes.

A checklist

  • Protocol written before collection, with sampling, window, vantage point, tooling, access rules and failure handling.
  • Ethics application covering expectations, identifiability, harm, minimisation and publication.
  • Raw responses preserved where the study allows, and every processing step counted.
  • Manifest with file hashes and collection context, README, licence and CITATION.cff.
  • Deposit with a persistent identifier, and a versioned dataset citation in the paper.

The same discipline applies outside academia; building large-scale training datasets covers it for machine learning corpora.

The bottom line

Scraped data can be rigorous research material, but only if it is treated like one: collected to a written protocol, reviewed for the ethics of its context rather than its public label, preserved with enough detail to explain what was seen, and published as a citable, versioned output.

Most of that is documentation and packaging, done once at the start and once at the end. It is the difference between a dataset that supports one paper and one that supports a field.

Sources and references

  • Association of Internet Researchers, Internet Research: Ethical Guidelines 3.0, 2019.
  • Wilkinson et al., The FAIR Guiding Principles for scientific data management and stewardship, Scientific Data, 2016.
  • Fiesler, Beard and Keegan, No Robots, Spiders, or Scrapers: Legal and Ethical Regulation of Data Collection Methods in Social Media Terms of Service, ICWSM 2020.
  • Citation File Format, version 1.2.0, and the cffconvert validator.
  • Zenodo, operated by CERN.
  • Code tested by Shifter on 6 October 2026.

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started