How to scrape responsibly ยท what each rule buys you in law

About this documentUpdated ShowHide

Sean McDermott, Co-Founder and CEO, UnGovr

Written by Sean McDermott (with AI assistance) using the LexLint law library, which supplied every legal instrument, status and date on these pages, and the handbook and insight documents on lexlint.io that carry the depth behind each one.

Every law named here links to its summary page on lexlint.io, translated to English (if needed) and restructured to a standard format for human and code use. Every case links to the court's or the regulator's own record where one could be reached.

© 2026 UnGovr, publishing as LexLint. The text and the figures are licensed under Creative Commons Attribution-ShareAlike 4.0: share and adapt them, including commercially, with credit to LexLint (UnGovr) and under the same licence. Please contact LexLint at hello@ungovr.org to discuss other terms. Logos and wordmarks belong to their owners.

Corpus figures as of .

Legal information, not legal advice. This document describes the law as written and dated; it does not apply it to any system. The notice at the foot says what that means.

Ask first. Identify yourself. Slow down. Keep a record. Do not take what is clearly not public, and do not pass on what you have no right to. Each is good manners, and each one closes a legal door.

1Seven rules, and what each one buys you

These rules are the code of conduct of the Library Carpentry lesson on the ethics and legality of web scraping, taught at the University of California, Santa Barbara (the lesson). The lesson says what to do. This document says what each rule buys you in law, which is the half the lesson does not.

None of the seven clears a project. Each one closes off a particular claim, or keeps a particular fact from counting against you. What is left depends on the places and the bodies of law the rest of this brief describes.

The ruleWhat it buys you in law
Ask first A yes is permission, in writing. A no is a refusal you now know about. A refusal withdraws whatever welcome you had, whether it comes as the answer to your question or as a letter later on, so asking does not put you in a worse place. It tells you the position you are actually in, sooner and more cheaply than a cease-and-desist letter would.
Identify yourself A user agent that names you, and says how to reach you, makes a refusal addressable to you by name. A refusal addressed to you by name is the strongest rung on the ladder in the access-controls document, so you want it early, while it costs little. You want to be easy to refuse, and just as easy to allow: a site can say yes to you by name as well as no.
Slow down A rate limit you keep under is what keeps trespass to chattels off the table. That old claim, for interfering with someone else's property, turns on harm to the site's servers, and a scraper that never strains them gives it little to stand on. It also keeps you from being the outage the site sues over.
Keep a record What you fetched, when, and under which robots.txt file and which terms is what you can be made to produce if a claim is brought. A record made at the time is also your own account of what you did. The evidence document in LexLint's handbook on AI agents says which records the law can ask for.
Do not take what is clearly not public A login wall, a paywall and a private area are each a contract, a lock or both. Getting past one moves you out of the law that helps with public pages and into contract law or computer-misuse law. Each is a rung on the ladder in the access-controls document.
Do not pass on what you have no right to share Copying for yourself and republishing are two different acts under copyright, and the second needs a right the first may not. An exception that covers the copy you keep, such as text and data mining or fair use, may not cover handing it on. The same goes for personal data: passing it on is a use of its own, which privacy law weighs separately, as the document on the people in the data explains.
Share what you can Data in the public domain, or data you have permission to share, published well with its source and its licence, is how the next person asks you first instead of scraping the same site again. It is the first rule, seen from the other side.

2The record to keep

Keep one record for each site, written at the time rather than rebuilt later. A short list is enough:

  • the site, and the date you began, with each date you went back;
  • the robots.txt file as you read it, saved as a copy, because sites change it;
  • the terms page as you read it, with its governing-law clause, the sentence that names whose law and whose courts the site chose;
  • the legal notice, the page that says who runs the site, and the operator it names;
  • your user agent string and the rate you ran at;
  • anything the site told you, and when: an email, a block, a letter;
  • the walkthrough's result.

The terms page and the legal notice are also two rungs of the ladder in the document on whose law applies, the one you climb to find where the site has legal standing, so the same copy does two jobs.

3What to take to a lawyer

This brief describes the law. It does not apply it to your project; that is a lawyer's work, and it goes further when the lawyer starts from the facts instead of spending the first hour collecting them. Take four things:

  • your result from the walkthrough: the places, the concerns in the order to deal with them, and the profile;
  • the record above;
  • the target site's terms and legal notice, as you read them;
  • the scraping sheet, two pages made to be printed and handed over.

The questions worth asking first are the concerns at the top of the walkthrough's list.

The long read. the evidence document carries the depth behind this document, with every citation and its date.