Part 3: Doing it

So you want to build a scraper ยท the options, with the legal angle at each fork

Updated Legal information, not legal adviceAbout this documentShowHide

Sean McDermott, Co-Founder and CEO, UnGovr

Written by Sean McDermott (with AI assistance) using the LexLint law library, which supplied every legal instrument, status and date on these pages, and the handbook and insight documents on lexlint.io that carry the depth behind each one.

Every law named here links to its summary page on lexlint.io, translated to English (if needed) and restructured to a standard format for human and code use. Every case links to the court's or the regulator's own record where one could be reached.

© 2026 UnGovr, publishing as LexLint. The text and the figures are licensed under Creative Commons Attribution-ShareAlike 4.0: share and adapt them, including commercially, with credit to LexLint (UnGovr) and under the same licence. Please contact LexLint at hello@ungovr.org to discuss other terms. Logos and wordmarks belong to their owners.

Corpus figures as of .

This document describes the law as written and dated; it does not apply it to any system. The notice at the foot says what that means.

Before you write a scraper, ask whether you need one: a bulk download, a feed or a documented API is a permission the scraping question never gives. If you do scrape, the forks after that are the API before the HTML, a service or your own, and, for your own, the tools and the rules.

This document is the route through the rest of the brief for someone about to build: the forks in the order you meet them, and at each one what the route gives you in law and what it leaves with you. The forks are the same for a first project written with a coding assistant and for an engineering team; only the tooling changes.

1Do you need to scrape at all?

Before you write a scraper, look for the ways the site already hands the data over. Each one is a permission, and the scraping question never gives you one: a page anyone can see is a fact about how you got the data, not a permission to have it, as the public-data document explains. In the order of how much permission each carries:

  1. A data licence, or a reseller. Some sites sell their data, or license it through a partner. A paid licence is the plainest permission there is, and it is often cheaper than the engineering it replaces.
  2. A bulk download, or an open-data portal. A file or a dataset the publisher offers whole, under a licence. The licence is the permission, and it says what you may do: credit the source, share alike, no commercial use, no model training. Read it before you download, because what governs a dataset is its licence, not how easy it is to fetch.
  3. A documented API. The site's own statement of what it will hand over, to whom, at what rate, and for what use. Registering for a key is accepting a contract, the clickwrap rung in the access-controls document, so the API's rate limit and its use restrictions are terms, and going over them is a breach rather than a technical foul. An API is also the best evidence there is of what the site welcomes.
  4. A feed, or a sitemap. An RSS or Atom feed is content offered for a reader to fetch on a schedule. A sitemap is the list of pages the site wants found. Both are invitations to fetch. Neither is a licence to reuse what you fetched; the terms page and copyright still decide that.

If one of these exists, most of the legal question collapses into one you can answer yourself: did I keep to the licence, or to the terms.

2Look for the API or the JSON before the HTML

Most pages today are built in the browser from data it fetches separately. Open the browser's developer tools, load the page with the network panel open, and filter for JSON or GraphQL: you will often find the endpoint the page itself calls, with the data already structured.

In law it is the site's, exactly as the page is, and the terms apply to it as they apply to the HTML. What matters is how the page reaches it. An endpoint the page calls with no login and no token tied to a user is the same public page in another format, and everything the public-data document says holds. An endpoint the page calls with a token it was given after a login, or one that answers only to the site's own app, is a login wall or a lock, and reproducing the token or the app's signature is stepping around it: the rungs in the access-controls document. Your user agent names you on a call to an endpoint as it does on a page, and an undocumented endpoint can change or close without notice, which is one more reason to keep the record that the careful-scraper document describes.

3A service, or your own?

A scraping service, whether a hosted crawler, a scraping API or a dataset built to order, sells the fetching. Its own terms will say the rest is yours, so it is worth knowing exactly what the fetching buys.

What the service carriesWhat stays yours
The machines, the addresses and the rendering. Where the scraper runs, the second of the three places in the document on whose law applies, moves to the service's place. The use. Copyright and the database right in what was collected, and every question about the people in the data: you decided what to collect and what to do with it, which is what makes you the controller under privacy law.
Its own crawl posture: its user agent, its rate, its reading of robots.txt, chosen by the service and written into its terms. The record. What was fetched, when, and under which terms is still what you can be asked to produce, and the service's logs are not yours unless the contract says so.
A contract with you. Read the indemnity clause first: it is the paragraph in which the service hands the legal exposure back. The counterparty. A site that objects writes to the party that wanted the data, and a cease-and-desist letter addressed to you is the strongest rung on the ladder however the pages were fetched.

Two things a service's brochure may offer are worth naming. Proxies are not licences: a proxy or a residential address changes which machine and which place a request comes from, and nothing about whether the site allows it. And a service that advertises solving CAPTCHAs or getting past bot challenges is selling the circumvention rung, the one that moves a scraper into computer-misuse law in the places that have it; that the service did it for you does not make the act someone else's. A service makes sense when you need scale or rendering you cannot run and the legal question about the data is settled.

4Your own

If you write it yourself, three documents in this brief carry the rest, and a fourth thing is a person.

  • The tools document names the standard kit for each job, in Python and JavaScript, and the setting at each row that matters in law.
  • The careful-scraper document's seven rules are the code of conduct, with what each one buys you and the record to keep.
  • The Scraping Checkup turns twelve answers into your scraping profile: the places, the concerns in the order to deal with them, and a lexlint.yml that the LexLint plugin reads, so a coding assistant that writes the scraper is told which law it is writing into. The FAQ says how to install it.

The plain limit. LexLint reads published law and dated cases, and it can say which bodies of law a scraper like yours walks into and what each one asks. It cannot read the target site's terms for you, it does not know what you will do with the data next year, and it has not read the cases that have not been published. It finds the obvious problems and it cannot find them all. For a real project, one with money in it, people in the data, or a counterparty who might mind, the next step is a lawyer who knows this area, scraping and data law rather than a general practice, because the answers turn on cases and instruments a generalist will not have read. Take them the record, the profile and the sheet; the careful-scraper document lists what to bring.

The long read. By continuing, you agree: when terms of use bind a crawler carries the depth behind this document, with every citation and its date.