Can I scrape public data? · public is not the same as permitted

About this documentUpdated ShowHide

Sean McDermott, Co-Founder and CEO, UnGovr

Written by Sean McDermott (with AI assistance) using the LexLint law library, which supplied every legal instrument, status and date on these pages, and the handbook and insight documents on lexlint.io that carry the depth behind each one.

Every law named here links to its summary page on lexlint.io, translated to English (if needed) and restructured to a standard format for human and code use. Every case links to the court's or the regulator's own record where one could be reached.

© 2026 UnGovr, publishing as LexLint. The text and the figures are licensed under Creative Commons Attribution-ShareAlike 4.0: share and adapt them, including commercially, with credit to LexLint (UnGovr) and under the same licence. Please contact LexLint at hello@ungovr.org to discuss other terms. Logos and wordmarks belong to their owners.

Corpus figures as of .

Legal information, not legal advice. This document describes the law as written and dated; it does not apply it to any system. The notice at the foot says what that means.

A page anyone can see is a fact about how you got the data, not a permission to have it. Privacy law, copyright and computer-misuse law each answer the question differently, and only one of them is on your side.

1Public is a fact about how you got it

The most expensive misunderstanding about scraping is that data anyone can see is data anyone can take. A page open to everyone is a fact about how you obtained what is on it. It is not a permission to have it, or to do what you like with it.

Three bodies of law each ask whether the data was public, and each means something different by the word. Only one of them makes public the whole question.

2Three bodies of law, three meanings

Body of lawWhat “public” means to itWhose side it is on
Privacy law Public personal data is still personal data in the EU and the UK. That it was public is a fact that goes into the balance, not around it, and the EDPB's 2026 guidelines on scraping, still open for comment, weigh it as one fact among several. California goes the other way: its Consumer Privacy Act takes information that is publicly available out of personal information (the publicly available information exemption), though it reads “publicly available” narrowly. The people in the data, in the EU and the UK, whom a later document takes up.
Copyright Public does not mean unprotected. A page anyone can read is still somebody's work, and copying it is what copyright reaches. The owner of the work, unless an exception applies.
Computer-misuse law Public is where the gate is open. The US Supreme Court's Van Buren decision (2021) asked whether a gate was closed, and hiQ v. LinkedIn (2022), in the Ninth Circuit and at the injunction stage, applied that to scraping: reading a page the site shows to everyone is probably not access without authorisation under the Computer Fraud and Abuse Act. That is one appeals court's reading, not a settled national rule, and hiQ itself lost on its contract later that year. Yours. The one body of law where public settles it.

In Germany, and in the rule the EU sets for all its member states, computer-misuse law also turns on getting past a protection, and a page with none has nothing to get past. The United Kingdom's is phrased differently and has not been tested against a crawler in a reported case.

Even in the United States, public answers one question only. hiQ won its computer-misuse argument in 2022, then lost before the year was out on the contract it had accepted when it opened an account, and agreed to a court order to stop scraping and destroy what it had built. Next, the access-controls document takes up logins and terms.

Copyright reaches the making of a copy, and a scraper makes a copy of every page it keeps. So the question is never whether the page was public. It is whether an exception to copyright covers what you do with the copy, and that depends on where the copy is made and what it is for.

In the United States the exception is fair use, a balance of the purpose, the kind of work, how much you take and the effect on the market for the original. Countries that follow the British model have fair dealing instead, a narrower list of named purposes such as research, criticism and news reporting. The EU has two exceptions for text and data mining: one for research bodies, and a general one, commercial use included, in Article 4 of the 2019 copyright directive. The general one covers only works you can lawfully reach, and it gives way wherever the owner has reserved their rights in a form a machine can read.

The locks on the door have their own law, and their own document.

4A compiled database has a right of its own

In the EU and a few other places, a database right protects the investment someone made in gathering, checking or presenting a collection of data, whatever the copyright in the items inside it. Taking a substantial part of the collection is a claim by itself, even when no single item you took was protected on its own. France's version is the database producer's right in its intellectual property code.

5Competing with the site using its own work

A last group of claims looks at what you build. Hot news, unfair competition and misappropriation reach a scraper that uses a site's own compiled work to offer something that competes with the site. In the United States hot news began in 1918, when the Supreme Court stopped one news wire copying another's freshly gathered war news, and the federal Copyright Act now overrides most such claims, leaving a narrow one (its preemption section). When a news agency sued a media-monitoring service that scraped its stories, it won on copyright in Associated Press v. Meltwater, and the court counted against the service that its excerpts stood in for the agency's own news sites. In some other countries a general unfair-competition law reaches scraping that copyright does not.

The long read. There is no robots.txt for people carries the depth behind this document, with every citation and its date.