Publicly available does not mean free to use
Generative AI development depends on data at scale. Web scraping provides a way to collect large datasets from publicly accessible sources including websites, social networks, forums, news outlets, public registers and open-data portals.
But public accessibility is not the same as unrestricted legal use.
In July 2026, the European Data Protection Board adopted Guidelines 03/2026 on web scraping in the context of generative AI. Version 1.0 was adopted for public consultation on 7 July 2026 and is not yet final. The consultation closes on 30 October 2026.
Their scope is focused but commercially important. They address scraping personal data from external internet sources by private organisations for the training or fine-tuning of generative AI models. They cover both organisations that scrape data themselves or through a contractor and organisations that obtain an already-scraped dataset for AI development.
Where the activity involves personal data, the GDPR applies to the relevant collection, extraction, cleaning, organisation, storage and subsequent use of that data.
The practical question is therefore not simply:
“Can the data technically be collected?”
It is:
“What data is being collected, for what purpose, from which sources, under which lawful basis, and with which safeguards?”
Responsibility has to be mapped before the dataset is built
Web scraping can involve several organisations: the AI developer, a specialist scraping provider, a dataset supplier or other intermediaries.
The EDPB stresses that GDPR roles depend on the actual circumstances rather than contractual labels.
Where an AI developer instructs another organisation to create a training dataset on its behalf, including determining relevant sources and data categories, the developer will generally determine the purposes and essential means of the processing and may therefore act as controller, with the scraping provider potentially acting as processor.
The position can be different where an organisation acquires a dataset that was already scraped independently. The scraper and AI developer may then be separate controllers for their respective processing activities.
Joint controllership may also arise where organisations jointly determine the purposes and essential means.
For AI development projects, this makes data-supply architecture part of the GDPR analysis.
Before acquiring or building a scraped dataset, organisations should know:
- who determines the collection criteria;
- who selects the sources;
- who decides which data categories are retained;
- who performs filtering and cleaning;
- who determines the AI-development purpose; and
- which organisation is responsible for responding to individuals.
A dataset contract cannot substitute for that factual analysis.
Purpose has to be defined before the data is collected
Purpose limitation applies independently of the lawful-basis analysis. Personal data must be collected for specified, explicit and legitimate purposes and must not subsequently be processed in a manner incompatible with those purposes.
For generative AI development, the purpose should therefore be defined before collection criteria and source-selection rules are established. Where the precise downstream use of a general-purpose AI model has not yet been determined, the EDPB recommends identifying the objective pursued by the development, including whether it is commercial, public or scientific in nature and whether the intended use is internal or external to the organisation.
A generic reference to AI development should therefore not replace a sufficiently defined processing purpose.
Legitimate interests may be available, but it is not a shortcut
For private-sector web scraping in generative AI development, the EDPB identifies legitimate interests under Article 6(1)(f) GDPR as a lawful basis that is often relied upon.
That does not mean that AI development automatically constitutes a legitimate-interest justification.
The full three-stage test remains necessary:
01 · LEGITIMATE INTEREST
The interest pursued must be lawful, clearly and precisely articulated, and real and present rather than speculative.
02 · NECESSITY
The processing must be necessary for that interest. Organisations must consider whether an equally effective but less intrusive means could achieve the purpose.
03 · BALANCING
The controller’s interests must be balanced against the interests, rights and freedoms of the individuals whose data is being processed.
The EDPB specifically links necessity to the design of the collection process. Narrower scraping criteria, pseudonymised data or synthetic data may in some circumstances provide less intrusive alternatives to broad collection.
This means that a legitimate-interest assessment cannot sensibly be conducted in isolation from the technical design of the scraping operation.
Putting information online is not GDPR consent
The Guidelines are particularly clear on one point.
Where an individual has made personal data publicly accessible online, this does not mean that the person has consented to that data being scraped for a particular AI-development purpose.
Similarly, the absence or non-applicability of a robots.txt file does not constitute consent within the meaning of the GDPR.
Consent will in many large-scale scraping scenarios be difficult to use as a lawful basis because there is generally no direct relationship with the individuals concerned and obtaining valid consent from every person before collection may be impracticable.
The fact that data is accessible therefore does not remove the need to identify another appropriate legal basis.
Reasonable expectations depend on context
Public accessibility is nevertheless relevant to the legitimate-interest balancing exercise.
The EDPB recognises that people may now understand that material published online can be accessed and reused by third parties. But this does not mean that individuals can reasonably be expected to anticipate every type of reuse, by every controller, for every purpose.
The assessment is contextual.
Relevant factors include:
- the nature of the website or platform;
- whether the material is genuinely freely accessible or subject to access restrictions;
- whether the individual made the information public themselves;
- the relationship, if any, between the individual and the organisation scraping the data;
- the nature of the personal data;
- the characteristics of the individuals concerned, including whether children or other vulnerable individuals are involved; and
- technical or other measures signalling that automated scraping is not permitted.
The EDPB specifically refers to measures such as robots.txt, ai.txt, CAPTCHA and access restrictions.
These signals do not themselves determine the GDPR lawful basis. But they can be relevant both to individuals’ reasonable expectations and to how an organisation designs its source-selection controls.
Data minimisation starts before the crawler runs
One of the strongest practical messages in the Guidelines concerns data minimisation.
Large datasets are not prohibited simply because they are large. But the scale of an AI training dataset does not remove the requirement that personal data be adequate, relevant and limited to what is necessary.
The EDPB therefore expects minimisation to affect the collection architecture itself.
Before scraping begins, organisations should consider measures such as:
- determining whether synthetic data could be used instead of personal data;
- defining precise collection criteria;
- mapping the types of information expected to be collected;
- excluding unnecessary categories of personal data;
- excluding sources structurally likely to contain information concerning children, vulnerable individuals or particularly sensitive information; and
- excluding websites that clearly oppose automated scraping through relevant technical measures.
During and after collection, controls may include:
- syntax-based filtering of identifiers;
- removing unnecessary information;
- replacing real data with synthetic data where feasible;
- anonymisation;
- pseudonymisation; and
- reviewing whether retained information continues to be necessary.
The compliance analysis therefore begins before collection, rather than after a raw dataset has already been assembled.
Transparency cannot simply be assumed away
Large-scale web scraping creates an obvious difficulty: the organisation may not know the identity or contact details of every person whose information appears in the dataset.
Article 14(5)(b) GDPR may, depending on the particular circumstances, allow an organisation not to provide information individually where doing so is impossible or would involve disproportionate effort.
But the EDPB warns against treating that exception as automatic.
The organisation must assess the circumstances, including the size and age of the dataset, the number of individuals concerned and the safeguards implemented.
And even where individual notification is not required, transparency does not disappear.
The information must still be made publicly available.
The EDPB’s proposed level of transparency is notable. It recommends, where possible:
- identifying the sources of the data with precision;
- explaining whether the sources are publicly or non-publicly accessible;
- describing relevant characteristics of the crawler;
- identifying scraped domains and URLs in searchable form where practicable;
- indicating the date or collection period;
- explaining how individuals can exercise their GDPR rights; and
- where an existing scraped dataset has been acquired, providing relevant information about the organisation from which it was obtained.
A generic privacy notice saying merely that “publicly available information may be used” may therefore fall well short of the governance model contemplated by the Guidelines.
Accuracy also matters at dataset and model level
Scraped information may be old, incomplete, duplicated or unreliable.
The EDPB recommends that organisations, to the greatest extent possible:
- scrape from reliable and maintained sources;
- record when information was collected; and
- validate data before it is used in AI training.
The Guidelines also connect accuracy with the resulting model. Where the model is expected to generate outputs that themselves constitute personal data, organisations should consider the risk that inaccurate training information contributes to inaccurate outputs.
Source governance is therefore relevant not only to the composition of the dataset but also to downstream AI behaviour.
Special-category data raises a substantially higher compliance challenge
Large-scale scraping can capture information revealing health, political opinions, racial or ethnic origin, sexual orientation or other special categories of personal data even where the organisation did not specifically intend to collect it.
Where an organisation intentionally processes special-category data, an Article 6 lawful basis is not enough. A condition under Article 9(2) GDPR is also required.
The more difficult issue is incidental and residual collection.
The EDPB considers that the CJEU’s reasoning in GC and Others (C-136/17) may be relevant in particular circumstances. But it expressly warns that this is not a general exemption from Articles 9 and 10 GDPR.
A case-by-case assessment is required.
The EDPB’s proposed application of that reasoning is narrow. It requires relevant similarities with the processing considered in GC and Others; only incidental and residual, rather than intentional, processing of special-category data; difficulty or impossibility in determining and preventing that collection in advance; and measures within the controller’s responsibilities, powers and capabilities to prevent the dissemination of such data.
Where the reasoning is relevant, the EDPB expects extensive measures within the controller’s responsibilities, powers and capabilities. These extend across the AI-development lifecycle.
Before collection, controls may include precise source criteria and filtering designed to prevent collection of special-category data.
After collection, relevant information that has nevertheless been captured should be identified and deleted where required.
During model development, organisations should consider controls against extraction, privacy attacks and disclosure through model outputs.
After development, output monitoring and reinforced filtering may be required. The Guidelines also point to model-unlearning or equivalent techniques as technologies that may become relevant as the state of the art develops.
This moves the Article 9 analysis beyond the dataset itself and into the design, testing and operation of the model.
Safeguards can determine whether legitimate interests is available
The Guidelines provide a useful contrast.
In one example, an organisation collects large volumes of publicly available voice recordings to develop a voice-generation tool without additional measures to protect the training data or reduce the risk of unlawful or malicious reuse. The EDPB considers that legitimate interests cannot support the processing in that scenario.
In another example, a developer of a text-generative AI system uses freely and publicly accessible material and implements measures addressing memorisation and regurgitation, problematic outputs, data-subject rights and source transparency. The EDPB considers that the balancing test may generally be satisfied in those circumstances.
The point is significant.
Safeguards are not merely measures added after a lawful basis has already been selected. They can be central to whether the balancing test succeeds at all.
Data-subject rights need to work in practice
The EDPB also identifies mechanisms supporting individual control as possible mitigating measures.
These can include making information about collection widely available, facilitating the exercise of GDPR rights and potentially offering mechanisms allowing individuals to object to collection or to place relevant identifiers on an opt-out list.
The technical challenge may be substantial where datasets contain billions of items or where identifying particular individuals is difficult.
But scale does not eliminate accountability.
Organisations should consider how requests relating to source data, retained datasets and subsequent model behaviour will be investigated and implemented before the system is deployed.
Scraped datasets require due diligence too
An organisation does not avoid these issues by purchasing a dataset instead of running the crawler itself.
The Guidelines expressly cover AI developers obtaining datasets that have already been scraped by another organisation.
The receiving organisation therefore needs to understand, among other things:
- who collected the data;
- which sources were used;
- what collection criteria applied;
- when the data was collected;
- which categories of personal data may be present;
- what filtering was undertaken;
- which lawful basis supported the relevant processing;
- how transparency was addressed;
- whether restrictions or objections to scraping were respected; and
- how individuals’ rights can be operationalised.
“Third-party dataset” should therefore not be treated as a provenance answer.
It is the beginning of the due-diligence inquiry.
GDPR is not the only source-governance question
The EDPB Guidelines are concerned with data protection, but web-scraping projects can also engage other legal regimes.
Notably, one of the EDPB’s own examples concerning a potentially sustainable legitimate-interest analysis assumes controls concerning copyright and text-and-data-mining rights alongside GDPR safeguards.
AI dataset governance may therefore need to address privacy, intellectual-property restrictions, platform conditions and other applicable regulatory requirements separately.
Compliance under one regime does not itself resolve another.
What organisations should check now
01 · PURPOSE
What precisely is the dataset being created or acquired to achieve?
02 · DATA AND SOURCES
Which personal data and source categories will the crawler encounter?
03 · ROLES
Who is controller, processor, separate controller or potentially joint controller?
04 · LAWFUL BASIS
If legitimate interests is relied upon, is the interest specific, is the processing necessary, and has the balancing test been documented?
05 · SOURCE SELECTION
Are access restrictions, robots.txt, ai.txt, CAPTCHA or other anti-scraping signals reflected in collection rules?
06 · MINIMISATION
What is excluded before collection, filtered during collection and deleted or transformed afterwards?
07 · TRANSPARENCY
Can the organisation explain the sources, collection period, crawler and rights mechanisms with sufficient specificity?
08 · SPECIAL-CATEGORY DATA
How will intentional, incidental and residual Article 9 data be identified and controlled?
09 · MODEL-LEVEL CONTROLS
What measures address memorisation, regurgitation, extraction, privacy attacks and problematic outputs?
10 · DATASET DUE DILIGENCE
If the data comes from a third party, can the organisation demonstrate where it came from and how it was collected?
11 · ONGOING GOVERNANCE
Are source lists, filtering controls, rights mechanisms and technical safeguards reviewed as the dataset and model evolve?
The compliance architecture begins upstream
The EDPB’s draft Guidelines move the web-scraping discussion away from the assumption that publicly available information is simply an available AI resource.
For generative AI development, GDPR compliance starts upstream.
It starts with the purpose of the model, the decision to use personal data, the selection of sources, the instructions given to the crawler and the categories of information that are deliberately excluded.
It continues through dataset cleaning, lawful-basis analysis, transparency, rights management and controls against sensitive-data leakage, memorisation and regurgitation.
The practical message remains the same as the original Privacy Minders analysis:
“Available online” is not a GDPR compliance strategy.
The more important question is whether the dataset has been designed, sourced and governed in a way that the organisation can actually defend.
Originally discussed by Privacy Minders, now OSTRAI, following publication of EDPB Guidelines 03/2026. This Regulatory Update expands the original LinkedIn analysis. Original LinkedIn post → EDPB Guidelines 03/2026 →



