Can I scrape names, emails and photos? · the people in the data

About this documentUpdated ShowHide

Sean McDermott, Co-Founder and CEO, UnGovr

Written by Sean McDermott (with AI assistance) using the LexLint law library, which supplied every legal instrument, status and date on these pages, and the handbook and insight documents on lexlint.io that carry the depth behind each one.

Every law named here links to its summary page on lexlint.io, translated to English (if needed) and restructured to a standard format for human and code use. Every case links to the court's or the regulator's own record where one could be reached.

© 2026 UnGovr, publishing as LexLint. The text and the figures are licensed under Creative Commons Attribution-ShareAlike 4.0: share and adapt them, including commercially, with credit to LexLint (UnGovr) and under the same licence. Please contact LexLint at hello@ungovr.org to discuss other terms. Logos and wordmarks belong to their owners.

Corpus figures as of .

Legal information, not legal advice. This document describes the law as written and dated; it does not apply it to any system. The notice at the foot says what that means.

Consent is not available to a scraper. Legitimate interest is the basis with a test attached. The duty to tell people does not go away because you cannot tell them one by one. And “they put it online” is an argument about one limb of the test, not a way around it.

1The park is public and the people are not

A page can be crawled and mined within the rules of copyright and computer-misuse law and still be unlawful to process, because some of what is on it is a person. Copyright asks who owns the words and the pictures; it has nothing to say about the people in them. In the EU and the UK, that a page was public is not a lawful basis for keeping the personal data on it. It is a fact about how you got the data, and it goes into the balance described below rather than around it.

2Consent is out, and legitimate interest has a test

Under the EU's General Data Protection Regulation, whoever keeps personal data needs a lawful basis for it. Consent is not available to a scraper, which has no relationship with the people on the pages and cannot ask them all. That is not a loophole: it leaves the harder basis. The EDPB said so in its Guidelines 03/2026 on scraping, adopted on and open for consultation until , so they are guidance that may still change.

Legitimate interest is the basis left, and it comes with a test in three parts: a real interest, clearly stated; necessity, meaning nothing less intrusive would do; and a balance against the people's own interests, weighing what the data is, where it was published and what they could reasonably expect. “They put it online” is an argument about the third part, not a way past the test.

The duty to tell people survives too. When you collect data about people from somewhere other than them, the law expects you to tell them. Being excused from telling each one, because you cannot, does not excuse telling them at all: you publish a notice anyone can find, and you offer a way to say no before you collect.

3The enforcement record is unusually clear

Most of the law on scraping is thin and recent; this corner is not. Regulators in several European countries have ruled against Clearview AI, which collected images of people from public web pages, and every one reached the same conclusion: the pages being public did not make the processing lawful. Their penalties total more than 100 million euros. The question of reach was answered separately. In October 2025 the United Kingdom's Upper Tribunal concluded that Clearview's processing related to monitoring the behaviour of people in the United Kingdom, so it fell within the country's data protection law although the company is based outside it. That is the ruling to notice if you work outside Europe.

4Where the regimes diverge

California goes the other way. Its Consumer Privacy Act takes information that is publicly available out of personal information altogether (the publicly available information exemption, since ), though it reads “publicly available” narrowly: government records, and what a person made public themselves or through widely distributed media, not a copy of it on a directory site they did not control. So one page can be treated two ways on the two sides of the Atlantic, and a scraper working on both sides is under both.

The long read. There is no robots.txt for people carries the depth behind this document, with every citation and its date.