Is web scraping legal? · there is no internet law, only the law of every place your scraper touches

About this documentUpdated ShowHide

Sean McDermott, Co-Founder and CEO, UnGovr

Written by Sean McDermott (with AI assistance) using the LexLint law library, which supplied every legal instrument, status and date on these pages, and the handbook and insight documents on lexlint.io that carry the depth behind each one.

Every law named here links to its summary page on lexlint.io, translated to English (if needed) and restructured to a standard format for human and code use. Every case links to the court's or the regulator's own record where one could be reached.

© 2026 UnGovr, publishing as LexLint. The text and the figures are licensed under Creative Commons Attribution-ShareAlike 4.0: share and adapt them, including commercially, with credit to LexLint (UnGovr) and under the same licence. Please contact LexLint at hello@ungovr.org to discuss other terms. Logos and wordmarks belong to their owners.

Corpus figures as of .

Legal information, not legal advice. This document describes the law as written and dated; it does not apply it to any system. The notice at the foot says what that means.

No body of law governs the internet. A scraper is under the law of every place it touches, and the places are found by asking where each party is: you, the machine, the site, the people on its pages, and the model if one is in the loop.

1There is no internet law

No body of law governs the internet as a whole. There is the law of each place: every country makes its own, and inside many countries each state or province adds more. A scraper is under the law of every place it touches, and it touches several at once: the place you work from, the place the program runs, the place the website belongs to, and the places where the people named on its pages live.

Three beliefs about scraping are common, and each is mistaken in a way that has a document of its own here. The first is that anything on a public page is yours to take; the public-data document explains why being able to see a page does not settle what you may do with what is on it. The second is that robots.txt is the law; the access-controls document explains what that file is and what it is not. The third is that the country where the website's server sits decides which law applies; the whose-law document, which comes next, explains why the server is almost never the answer.

2The parties around a scraper

In law, a party is a person or organisation that holds rights or owes duties in a situation, and so can be held to account. LexLint's handbook on software law names six parties around any program that acts online, and a scraper has the same six. Each one has a location, and each location can bring its law with it.

The six parties around a scraper. Each box opens that party's entry in the handbook's parties document.

MAKER is whoever wrote the scraper: you, if you or someone working for you built it, or the vendor or the AI model provider, if they did.

OPERATOR is whoever runs it: you, the cloud service it runs on, or a vendor that runs it for you.

COUNTERPARTY is the site the scraper reads, the company behind it and anyone whose work is published on it. It is the party this whole brief turns on, because most of the law that is particular to scraping exists to protect it.

USER, for a scraper, means the people on the pages: the names, faces, posts and contact details you collect. They are affected by what you do with their details, which are personal data, and each of them is a data subject with rights under the privacy law of the place they live.

OVERSEER means the regulators and courts of every place the other parties bring in, so more places mean more of them.

DISTRIBUTOR is rarely present for a scraper, because the role belongs to whoever sells or passes on software someone else made, and most scrapers are built and run by the same organisation.

3The questions that find the places

Four questions find the places whose law applies. The fourth matters only when an AI model reads what you collect or learns from it.

  1. Where are you?
  2. Where does the scraper run?
  3. Where does the site have legal standing? Often more than one place.
  4. If a model is in the loop: whose, and where is its provider?

The answer is nearly always more than one place. You know the first two already. The third takes work: a site can belong to a company in one country, sell in a second and name a third in its terms. Next, the whose-law document is how to find it.

4What scraping shares with the law of AI agents, and what is its own

Shared with the law of AI agentsScraping's own
the parties and where they are; the hook rule per body of law the target site as the central counterparty, where the handbook's centre is the user
privacy law over the people in the data computer-misuse and unauthorised-access law
AI law when a model collects, reads or is trained on the pages terms of use forming a contract with a machine
what a record has to show afterwards anti-circumvention law (DMCA §1201 and its equivalents)
cross-border transfer database rights, where they exist
text-and-data mining exceptions and the right to reserve against them
unfair competition, hot news, trespass to chattels
robots.txt and the other crawl signals
revocation: a cease-and-desist letter, an IP block, a changed terms page

The left column is ground this brief shares with LexLint's handbook on the law of AI agents, which covers it in depth, starting with the handbook's parties document. The right column is this brief's own: the law that exists because of what a scraper does to a site, taken up in the documents that follow.

The long read. Introduction: The 6 parties in AI law carries the depth behind this document, with every citation and its date.