For years, one of the fundamental problems in platform research has been asymmetry. Digital platforms can observe enormous amounts of activity across their services, while researchers outside those platforms often have to work with whatever data the platform chooses to expose.
The European Union’s Digital Services Act is beginning to change that relationship.
The Digital Services Act data access framework is moving from regulatory language into operational research infrastructure. On 19 May 2026, Digital Services Coordinators had already received 49 applications for vetted researcher status. According to the European Commission, those applications primarily concerned social media platforms and included research into illegal content, advertising transparency, and AI features.
That is significant. But access to platform-held datasets is only part of the research problem. If researchers want to understand how digital platforms affect people in Europe, they also need ways to observe those platforms from the user’s side of the interface. Official datasets can answer some questions. Independent observation can answer others.
Key takeaways
- The DSA is turning researcher access from a legal principle into an operational research regime.
- Platform-provided datasets and independent outside-in observation answer different, complementary questions.
- X’s July 2026 corrective measures reinforce that eligible researchers may need automated access to publicly available platform data.
- Geography, device, session state and time should be treated as research variables when studying what users actually experience.
- The next stage of platform accountability will increasingly depend on reproducible observations with clear provenance.
What Digital Services Act data access actually changes
The DSA establishes several important routes through which independent researchers can study very large online platforms and very large online search engines.
One concerns internal platform data. Under Article 40, vetted researchers can seek access to data needed to study systemic risks in the European Union and evaluate the effectiveness of platforms’ risk mitigation measures.
The process has become substantially more concrete. In July 2025, the European Commission adopted a delegated act setting out technical conditions and procedures for researcher access. It also created a DSA data access portal through which researchers, platforms, Digital Services Coordinators, and other relevant authorities can manage parts of the process.
Researchers seeking internal data need to meet important safeguards. These include affiliation with a research organization, independence from commercial interests, appropriate data security and confidentiality measures, protection of personal data, disclosure of research funding, and a commitment to publish their findings.
But the DSA also recognizes another category of evidence: publicly available platform data. The Commission explains that researchers meeting the relevant conditions can access publicly available data from very large online platforms and search engines for research into systemic risks without the intermediation of a Digital Services Coordinator.
That distinction matters. An internal dataset can tell researchers what a platform recorded. An API can tell them what a platform made available through that interface. A transparency repository can tell them what the platform disclosed. None of those necessarily tells a researcher exactly what a European user experienced on a particular screen, in a particular country, at a particular moment. That requires observation.
The X decision is more important than it first appears
The European Commission’s enforcement against X makes this distinction unusually visible. In December 2025, the Commission fined X 120 million euros for several DSA transparency breaches. Among them was a failure to provide researchers with appropriate access to public data. The Commission specifically said X’s terms prohibited eligible researchers from independently accessing public data, including through scraping, and that its processes created unnecessary barriers to research into systemic risks in the EU.
The next development is even more interesting from a research-methodology perspective. In July 2026, the Commission accepted an action plan under which X committed to improve and accelerate researcher screening, provide eligible researchers with access free of charge, provide appropriate volumes of data, and amend its terms so eligible researchers are not contractually prohibited from scraping publicly available data.
That should not be interpreted as a blanket legal authorization for unrestricted scraping. Researchers still need to operate within the applicable DSA conditions and other legal, ethical, privacy, and security requirements.
But it does establish an important principle. Independent collection of public information is not necessarily an activity sitting outside platform accountability. Under the right conditions, it can be part of the accountability infrastructure itself. The Commission has taken a similar position elsewhere. In its DSA commitments with AliExpress, the platform agreed to enable researchers who meet the Article 40(12) criteria to independently access and use public data to study systemic risks through automated methods such as data scraping.
The regulatory conversation is therefore moving beyond a simple question of whether researchers receive an API key. It is moving toward whether researchers can independently study what platforms actually do.
APIs tell you what the platform exposes, observation tells you what the user experiences
Official access remains essential. A platform-held dataset may contain moderation histories, internal classifications, aggregate metrics, system records, or historical information that would be impossible to reconstruct externally. Research APIs can make large-scale analysis easier. Advertising repositories can provide structured evidence about campaigns. Transparency databases can expose important information about moderation decisions and enforcement activity. These are forms of inside-out evidence.
But accountability also requires outside-in evidence.
Consider the kinds of research questions that increasingly matter. What does a user in Madrid see when they make a particular query? Is an advertisement visible in Italy but not the Netherlands? Are political or commercial disclosures presented consistently between markets? Does a recommendation interface return different outputs from different European countries? Does the mobile experience expose information that the desktop experience does not? Are platform changes deployed simultaneously across the EU? Can an observed behavior be reproduced a day, week, or month later?
These are not necessarily questions about what exists inside a database. They are questions about how a system behaves when someone interacts with it.
The strongest research designs will increasingly combine both perspectives. Platform-reported data tells researchers what the system records. Independent observation helps establish what the system presents. Neither is inherently superior. They measure different things. And when the two disagree, that disagreement can itself become an important research finding.
Geography is a research variable, not just a proxy setting
This is where internet infrastructure becomes part of research methodology. Web platforms are not necessarily uniform environments. Geography can affect availability, advertising, language, interfaces, search outputs, regulatory disclosures, product features, and other elements of an online experience.
So when a researcher says “this is what the platform displayed”, there is an obvious follow-up question: displayed to whom, and from where?
At Shifter, we think location should increasingly be treated as an experimental variable. Our residential proxy network provides access to residential IPs across 195+ countries, with geographic targeting that can be used to establish controlled internet vantage points. That infrastructure can support a researcher who needs to observe the same publicly accessible interface from several defined geographic locations rather than assuming that one connection represents the experience of every European user.
That does not mean a residential IP automatically reproduces every element of a human user’s experience. Platforms may consider many signals beyond IP geography, including cookies, account history, language settings, device characteristics, and session state. That is precisely why the methodology matters. Geography should be recorded alongside those other variables, not treated as an invisible property of the research infrastructure.
Regulatory-grade observation needs a method, not just access
Once outside-in web observation starts contributing to regulatory research, the standard of evidence has to rise with it. “We opened the page and saw this” is not a methodology.
A rigorous observation should be accompanied by enough contextual information for another researcher to understand, challenge, and ideally reproduce it. That record might include the observation timestamp; target URL or platform surface; exit country and city; ASN or ISP where relevant; device and browser profile; language settings; authentication status; account or session state; experimental condition; observed output; screenshot, HTML, or structured evidence where appropriate; repeat-test results; and validation status.
In effect, researchers need a chain of provenance for web observation. We have set out what that record should contain in the case for a vantage-point standard.
The same principle already shapes how we approach infrastructure benchmarking. We do not compare two networks using different request volumes, machines, or test conditions and then treat the result as meaningful. Tests use matched request counts, consistent concurrency and targets, and controlled environments. We also distinguish between IPs actually observed during the test window and the much larger advertised network totals that cannot be measured in one moment.
Research into platform behavior demands the same discipline. If location is changed, keep the other variables stable. If device type is changed, control geography. If a surprising result appears once, repeat it. If a difference persists across multiple observations, record the conditions under which it persists.
DSA accountability is becoming a measurement problem. Internal data can reveal what platforms record; outside-in observation can test what users actually experience. The strongest evidence will increasingly come from triangulating the two under controlled, reproducible conditions.
From scraping infrastructure to observation infrastructure
Proxy technology has historically been discussed in fairly operational terms. How many requests can you make? How quickly can you rotate IP addresses? How reliably can you collect information at scale?
Those questions still matter for many applications. Our Web Scraping API, for example, is designed to remove much of the infrastructure involved in collecting public web data, while our SERP API provides structured search results across geographic locations and devices.
But there is another way to think about proxy infrastructure. For researchers, the important capability may not simply be collection. It may be controlled presence.
A research system should be able to say that one observation represented a connection from France, another represented Germany, both were performed within a defined time window, and the remaining experimental conditions were kept as similar as possible.
That changes the question from “how much of the web can we collect?” to “how reliably can we reproduce the conditions under which the web was observed?” That is a very different infrastructure problem. And as the DSA research ecosystem matures, we believe it will become a more important one.
What a DSA research stack could look like
A mature platform-accountability project will rarely depend on one source of evidence. Instead, we expect research teams to build layered systems.
| Evidence layer | What it provides | Examples |
|---|---|---|
| Platform-provided evidence | Information held or disclosed by the platform | DSA datasets, official APIs, ad repositories, transparency disclosures |
| Controlled observation | Repeatable evidence from the user-facing service | Defined geographic vantage points, device profiles and sessions |
| Provenance metadata | Context needed to understand and reproduce an observation | Timestamp, country and city, ASN or ISP, device, language and session state |
| Comparison and validation | Tests whether a result is isolated or systematic | Cross-market comparisons, repeat tests and controlled variable changes |
| Research governance | Controls around responsible collection and use | Data minimization, security controls, ethical review and documented purpose |
Residential proxy infrastructure belongs primarily in the controlled-observation and provenance layers. It is not a substitute for DSA access. It does not replace internal datasets. And it does not answer every research question. Its role is to help establish and reproduce a known external vantage point.
That distinction is important because this should not become a debate about APIs versus scraping. It should become a debate about evidence quality.
The next transparency debate will be about reproducibility
The DSA is already changing what platforms must disclose. The next challenge is determining whether claims about platform behavior can be independently tested.
If a platform says a particular disclosure appears to European users, can a researcher reproduce it? If an advertising repository reports one thing, does the public interface show another? If researchers identify a recommendation pattern in one Member State, does it appear elsewhere? If an interface changes after regulatory scrutiny, can independent observers document when and where the change occurred?
And if a study is repeated six months later, has enough contextual metadata been preserved to understand whether differences came from the platform, the location, the device, or the research method?
Those questions move platform transparency away from simple disclosure and toward reproducibility. That may be one of the DSA’s most important long-term effects. It is not only creating mechanisms through which researchers can request data. It is helping create the conditions for a more rigorous science of platform observation.
The bottom line
Digital Services Act data access represents a major improvement in the ability of independent researchers to understand Europe’s largest online platforms. But platform-held data is only one perspective.
To understand how digital systems affect users, researchers also need to observe those systems from the outside: across countries, devices, sessions, interfaces, and time.
The future of platform accountability is therefore unlikely to be built around a single evidence source. It will depend on triangulation: what the platform records, what the platform discloses, and what independent researchers can reproduce.
At Shifter, we believe geographic vantage-point infrastructure can play a useful role in that third category.
If regulators and researchers want to understand how Europe’s largest platforms are experienced by European users, the location from which those platforms are observed is not a technical footnote. It is part of the evidence.
This article is provided for general information only and does not constitute legal advice.