Part 3: Doing it

Which tools should I use to scrape? ยท the standard kit, and the legal angle of each row

Updated Legal information, not legal adviceAbout this documentShowHide

Sean McDermott, Co-Founder and CEO, UnGovr

Written by Sean McDermott (with AI assistance) using the LexLint law library, which supplied every legal instrument, status and date on these pages, and the handbook and insight documents on lexlint.io that carry the depth behind each one.

Every claim this page makes about a tool was read from the tool's own documentation on .

Every law named here links to its summary page on lexlint.io, translated to English (if needed) and restructured to a standard format for human and code use. Every case links to the court's or the regulator's own record where one could be reached.

© 2026 UnGovr, publishing as LexLint. The text and the figures are licensed under Creative Commons Attribution-ShareAlike 4.0: share and adapt them, including commercially, with credit to LexLint (UnGovr) and under the same licence. Please contact LexLint at hello@ungovr.org to discuss other terms. Logos and wordmarks belong to their owners.

Corpus figures as of .

This document describes the law as written and dated; it does not apply it to any system. The notice at the foot says what that means.

The tools are open source, mature and free, and none of them decides how it is used. The standard kit in Python and JavaScript, by job, with the setting at each row that keeps you on the right side of a rule the rest of this brief explains, and the two things no tool should be asked to do.

The table names the standard kit for each job, in Python and in JavaScript, then says which setting or habit at that row keeps you on the right side of a rule the rest of this brief explains. Where a tool has a switch for it, the switch is named. Every claim about a tool was read from the tool's own documentation on the date in the About block above; tools move, so check the switch before you rely on it.

1The kit, by job

The jobPythonJavaScriptThe legal angle
Fetching a page requests, aiohttp fetch (built into Node.js since version 18), undici Set the User-Agent header to a string that names you and says how to reach you: the second of the careful-scraper document's rules, and the one that makes a refusal addressable. Never copy a browser's or another product's string to look like something you are not. Treat a 429 response as an instruction, not an obstacle: it may carry a Retry-After header, and waiting that long is the site's rate limit kept.
Reading the page Beautiful Soup, lxml, selectolax cheerio, linkedom A parser reads what you fetched; the legal weight is in what you keep. Extract the fields you need and drop the rest. When a person is in the data, keeping less is a duty privacy law calls data minimisation, as the document on the people in the data explains, and under copyright a table of facts is a smaller question than a copy of the article they came from.
Crawling many pages Scrapy Crawlee (also in Python) A framework's defaults are tuned for throughput, not manners, and the settings that matter in law are the ones you turn on. Scrapy obeys robots.txt only when ROBOTSTXT_OBEY is true (the project template sets it; the built-in fallback is false), waits between requests only by DOWNLOAD_DELAY (the template's one second; the fallback is none) or AutoThrottle, and sends its own name unless USER_AGENT is yours. Crawlee obeys robots.txt only with respectRobotsTxtFile on, and maxRequestsPerMinute is its brake. Colly, the Go equivalent, ignores robots.txt until IgnoreRobotsTxt is set to false, and its Limit rule sets the delay.
Pages that need a browser Playwright, Selenium Playwright, Puppeteer A browser runs the site's own code and looks like a visitor, which is fine for a page that needs script to render and changes no rule. It is not for getting a CAPTCHA or a bot challenge to pass, and not for logging in with an account you were not given: both are rungs in the access-controls document where a careful crawler stops. The forks of these tools built to hide that a browser is automated exist for that one purpose, and this page does not name them.
Reading robots.txt urllib.robotparser (standard library), Protego robots-parser A robots.txt file asks rather than locks, the first rung of the ladder, but it is the site's clearest statement of what it does not want fetched, and a scraper that ignored it carries that fact into any dispute. All three parsers read the Crawl-delay line too. Keep the copy you read, dated: sites change it.
Going slowly Scrapy's AutoThrottle, aiolimiter bottleneck, p-limit A rate you keep under is what keeps trespass to chattels, the old claim for interfering with someone else's property, off the table, and keeps you from being the outage the site sues over. One request at a time per site, with a pause between, is the default to start from; a Crawl-delay line or a 429 is the site saying its own number.
Keeping the record SQLite or PostgreSQL for the data, warcio for the fetch SQLite or PostgreSQL, warcio (the JavaScript port) Store what you fetched with when, from which address, and under which robots.txt and terms. A WARC file holds each response as received, headers and time included, which is the record the careful-scraper document says to keep and the evidence document in the handbook says the law can ask for.

2A coding assistant

For a first project, the tool is a coding assistant, and it will write any row of the table above on request. It will also write the circumvention code if asked, because it does not know which rule that walks into. The LexLint plugin is the row that tells it: the Scraping Checkup writes a lexlint.yml from your answers, naming the places and the concerns, and the plugin lints the project against the law of those places as the assistant works. The FAQ has the two lines that install it. The assistant is still the author, and you are still the party.

3Two things no tool should be asked to do

  • Solve a CAPTCHA, or get a bot challenge to pass. Each is a measure put there to keep automated visitors out, and getting past one is the circumvention rung: the one that moves a scraper into computer-misuse law in the places that have it. A service that does it for you does not move the act to the service.
  • Take a login you were not given. Credentials, a session token or a paid account's cookies are a contract someone else accepted and a lock someone else holds the key to. Behind them is authorised access law, the body that carries criminal exposure.

LexLint finds the obvious problems from published law and it cannot find them all: it does not read the site's terms for you, and it has not read the cases that are not yet published. For a real project, the next step is a lawyer who knows scraping and data law, with the record above in hand.

The long read. Disallow: / - what a website can actually forbid carries the depth behind this document, with every citation and its date.