Knowledge

Writing a Company Data Collection Policy: A Working Template

Most teams that collect web data have no written rules for it. A working policy template: scope, approvals, sources, personal data and takedowns.

Matt Brown

Matt Brown

October 4, 2026 · 9 min read

Most companies that collect web data have no written rules for doing it. A developer builds a scraper for a pricing project, another team copies it for lead research, a contractor adds a third, and nobody can say which sites are collected, what personal data is stored, or who would answer a complaint from a site owner. The work is usually fine. The problem is that nobody can show it is fine.

A short internal policy fixes that. It gives engineers clear defaults, gives legal and security teams something to review once instead of every time, and gives the company an answer when a customer, an auditor or a website asks how it collects data. This guide explains what such a policy needs to cover and gives a template you can adapt.

Key takeaways

  • A collection policy should be short enough that engineers read it: scope, an approval step, rules for sources, rules for personal data, and what happens when someone objects.
  • Make the safe path the default. Public, logged-out pages, honest identification, modest pace and robots.txt respected need no special approval; everything else does.
  • Treat objections as objections. France’s data protection authority, for one, expects collectors to exclude sites that object through robots.txt or CAPTCHAs.
  • Personal data changes everything: define what you may collect, minimise it at collection time, set retention, and make deletion possible.
  • Name an owner and a takedown process. The first complaint is the wrong time to decide who answers it.
  • This template is a starting point, not legal advice; have counsel review it against your jurisdictions and contracts.

Why a written policy matters

There are three practical reasons.

Consistency. Without written rules, each project makes its own decisions about robots.txt, pacing, personal data and retention. Some will be careful, some will not, and the company carries the risk of the least careful one.

Speed. A policy with clear defaults lets most projects start without a meeting. Legal and security review the policy once, then only the exceptions.

Evidence. Data protection regulators expect controllers to be able to show what safeguards they apply. The French data protection authority, the CNIL, for example, has published guidance on collecting personal data by web scraping that lists measures such as defining collection criteria in advance, excluding sites that clearly object, including through robots.txt or CAPTCHAs, filtering out unnecessary data, and deleting sensitive data as soon as it is identified. A written policy is how you show those measures exist.

What the policy needs to cover

SectionThe question it answers
Purpose and scopeWhich activities and teams does this apply to?
RolesWho owns the policy, who approves projects, who answers complaints?
Default rulesWhat can any project do without asking?
ApprovalWhat needs sign-off, from whom, with what information?
SourcesWhich sites and pages are in bounds, and how do we treat their signals?
Personal dataWhat may we collect about people, and how do we minimise and protect it?
Technical conductHow do collectors identify themselves, pace requests and handle credentials?
VendorsWhat do we require from proxy and data providers?
Storage and retentionWhere does collected data live, who can access it, how long do we keep it?
Objections and takedownsWhat happens when a site owner, a person or a regulator objects?
Records and reviewWhat do we log, and when do we review the policy?

The template

Adapt the wording to your company, delete what does not apply, and keep it short. Text in brackets is for you to fill in.

1. Purpose and scope

This policy governs the automated collection of data from websites and online services by [Company], its employees and its contractors, including scrapers, crawlers, browser automation, and data bought from third parties that collect on our behalf. It does not cover data our own users give us, or official APIs used under their own terms, except where section 6 applies.

2. Roles

  • Policy owner: [role, for example Head of Data] maintains this policy and the register of collection projects.
  • Approvers: [Legal contact] and [Security contact] approve projects that need approval under section 4.
  • Project owner: every collection project names one person responsible for following this policy.
  • Takedown contact: [role and shared mailbox] receives and answers objections under section 10.

3. Default rules for every project

Any project may proceed without further approval if it:

  • collects only publicly available pages that any visitor can see without logging in;
  • respects robots.txt and other machine-readable opt-outs that apply to the collector;
  • identifies the collector honestly and does not disguise automated traffic as a particular person;
  • paces requests so that it places no noticeable load on the target;
  • collects no personal data beyond what section 6 allows;
  • is recorded in the project register before it starts.

4. Projects that need approval

A project needs written approval from the approvers before it starts if it:

  • collects personal data beyond business contact details published for business purposes;
  • collects any special category data, such as health, religion or political opinion;
  • collects from pages behind a login, a paywall or any other access control;
  • continues after a site has shown it objects, including through robots.txt, a CAPTCHA, blocking, a contractual term or a direct request;
  • collects for training, fine-tuning or evaluating machine learning models;
  • sells, licenses or shares the collected data outside [Company].

The request states the purpose, the sources, the data fields, any personal data and its legal basis, the expected volume, the retention period and the project owner.

5. Sources and their signals

  • Prefer an official API, data feed or licence where one exists and covers the need.
  • Read and record the relevant terms of each source before collection starts.
  • Treat robots.txt, rate limits, CAPTCHAs, blocks and published reservations as signals of the site’s wishes, not as obstacles to engineer around.
  • Do not circumvent access controls, technical protection measures or authentication.
  • Stop collecting from a source promptly when it objects, and record the stop in the register.

6. Personal data

  • Collect only the personal data the stated purpose needs, and filter out other personal data at collection time where possible.
  • Never collect special category data unless approved under section 4; if it is collected by accident, delete it as soon as it is identified.
  • Pseudonymise identifiers where the analysis does not need the identity itself.
  • Do not combine collected data with other sources to identify individuals unless approved.
  • Make it possible to find and delete a person’s data on request.

7. Technical conduct

  • Store credentials for proxies, APIs and target accounts in the approved secrets manager, never in code.
  • Use only proxy and data providers that meet section 8.
  • Record, for each collection run, the time, the source, the collection location and the configuration used.
  • Monitor request volumes and error rates, and pause collection automatically when a source starts refusing requests.

8. Vendors

Proxy and data vendors must be able to show how their network or data is sourced and with what consent, operate a know-your-customer process, enforce an acceptable use policy, and sign terms consistent with this policy. The policy owner keeps a list of approved vendors.

9. Storage and retention

  • Store collected data only in [approved systems], with access limited to people who need it.
  • Keep raw collected pages for no longer than [period] and extracted data for no longer than [period], unless an approval sets a different period.
  • Delete data at the end of its retention period, and record the deletion.

10. Objections and takedowns

  • Any objection from a site owner, an individual or an authority goes to the takedown contact within [one working day].
  • The affected collection pauses while the objection is reviewed, unless legal advises otherwise.
  • The takedown contact responds within [period] and records the objection, the decision and any data deleted.
  • Repeated objections about the same project trigger a review of that project’s approval.

11. Records and review

The policy owner keeps the project register, approvals, vendor list and takedown log, and reviews this policy at least [annually] and whenever the law or the company’s activities change materially.

Making it stick

A policy that sits in a shared drive changes nothing. A few habits make it real:

  • Put the register where projects start. A short form in the tool engineers already use, with the default rules as checkboxes, catches projects before they exist rather than after.
  • Build the defaults into code. A shared collection library that respects robots.txt, sets an honest User-Agent, paces requests and records collection context makes the policy the easy path, not an extra step. Respecting robots.txt and AI opt-outs covers which signals to read.
  • Record where data was observed. Logging time, location and configuration per run is what makes collected data defensible later, as argued in the case for a vantage-point standard.
  • Check your vendors. Ask proxy providers how they source their IPs, and see how providers ethically source residential IPs and the malware economy behind cheap proxies for why it matters.
  • Keep secrets out of code. Running scrapers in CI/CD shows how to handle proxy credentials safely.
  • Read the signals before you build. A quick check of what protects a source, as in our anti-bot stack lookup, tells you early whether a project falls under the default rules or needs approval.

For the wider legal picture, see is web scraping legal and residential proxies and GDPR; for day-to-day practice, web scraping best practices.

The bottom line

A data collection policy does not need to be long. It needs a clear scope, safe defaults that let most work proceed, an approval step for the risky cases, firm rules for personal data, and a named person who answers when someone objects.

Write it once, build its defaults into the tools engineers use, and keep the register current. Then the next time someone asks how your company collects web data, the answer is a document rather than a scramble. Adapt the template to your situation, and have counsel review it before you adopt it.

Sources and references

Ready to get started?

Try Shifter's residential proxies, 205M+ IPs, 195+ countries, from $0.10/GB.

Get Started