Can I get past robots.txt, a CAPTCHA or a login? ยท what is in the way, and what it means to get past it
About this documentUpdated ShowHide
Sean McDermott, Co-Founder and CEO, UnGovr
Written by Sean McDermott (with AI assistance) using the LexLint law library, which supplied every legal instrument, status and date on these pages, and the handbook and insight documents on lexlint.io that carry the depth behind each one.
Every law named here links to its summary page on lexlint.io, translated to English (if needed) and restructured to a standard format for human and code use. Every case links to the court's or the regulator's own record where one could be reached.
© 2026 UnGovr, publishing as LexLint. The text and the figures are licensed under Creative Commons Attribution-ShareAlike 4.0: share and adapt them, including commercially, with credit to LexLint (UnGovr) and under the same licence. Please contact LexLint at hello@ungovr.org to discuss other terms. Logos and wordmarks belong to their owners.
Corpus figures as of .
Legal information, not legal advice. This document describes the law as written and dated; it does not apply it to any system. The notice at the foot says what that means.
A robots.txt file asks. A rate limit pushes back. A bot challenge and a CAPTCHA stand in the way. A login and a terms page make a contract. A letter revokes. Each rung moves you into a different body of law, and a court lit one rung in July 2026.
1Eight rungs, weakest first
The things a site can put between a scraper and a page come in eight kinds, and they are not equal. Read from the weakest to the strongest, each one moves you into a different body of law, so getting past each one means something different.
- A public page Nothing stands between the page and anyone who asks for it. In law the gate is open: under the US computer-misuse statute, on the Ninth Circuit's 2022 reading in hiQ, reading what a site shows everyone is not access without authorisation; not every court has said so. It moves you into no new body of law, and the public-data document covers the law that still applies to what is on the page.
- A robots.txt disallow A line in a file at the root of the site asking crawlers to stay out of some of its pages. In law it is a request, not a wall, and in the EU it is also a rights reservation against mining the site's pages. It moves you into copyright's rules on mining, and everywhere it is evidence of what the site wanted and what you knew.
- A rate limit or an IP block The site slows your requests down, or refuses the ones that come from your address. In law it is a technical measure of the weakest kind: it limits how fast you go, not where you go. Where your requests harm the server, it moves you toward trespass to chattels, an old claim for interfering with someone else's property.
- A bot challenge A test, set by the site or by a service in front of it such as a firewall or a content delivery network, that decides whether you are a person before the page is served. In law it is a measure that controls access, the kind a court treated as a lock in July 2026 (section 2 below). It moves you into computer-misuse law and, where the page is a copyrighted work, into anti-circumvention law.
- A CAPTCHA A puzzle that asks the visitor to prove they are a person. In law it is a challenge a program is not meant to pass, so a program that passes one has got around a measure on purpose. It moves you into the same law as the rung before, and a careful crawler stops here.
- A login wall Pages that show only after you sign in to an account. In law, signing in puts you inside a contract with the site, and the public-page rule that helped you on the first rung no longer does. It moves you into contract law as well as computer-misuse law.
- Terms you accepted The site's contract with its users, agreed by ticking a box or opening an account, or posted behind a link at the foot of every page. In law, terms you actively accepted (clickwrap) bind you, and terms behind a link nobody has to click (browsewrap) bind only where you had fair notice of them. It moves you into contract law, where a claim does not wait on computer-misuse law: hiQ, which won the 2022 public-data ruling, lost on its contract later that year.
- A cease-and-desist letter, or a block that follows one The site writes to tell you to stop, or blocks you after telling you. In law it is a revocation: whatever welcome you arguably had is withdrawn. It is the rung most bodies of law read the same way: what you do after it breaks the site's terms, and in many places it is what turns the next request into unauthorised access. The US computer-misuse statute is the exception on a public page: LinkedIn had written to hiQ and blocked it, and the Ninth Circuit still thought the claim unlikely, because the pages were public. Behind a login, or after a block on an account, the letter is what makes the next request unauthorised.
2The rung a court lit in July 2026
Two rulings from the same New York federal court, within eight months of each other, mark where the line now sits under the US anti-circumvention law, section 1201 of the Digital Millennium Copyright Act. The first was Ziff Davis v. OpenAI, in December 2025. The court dismissed an anti-circumvention claim built on a crawler ignoring robots.txt, because such a file controls access to a site no more than a sign asking visitors to keep off the grass controls access to a lawn.
The second was Reddit v. Perplexity, on . The same court let Reddit's main anti-circumvention claims go forward against Perplexity and a scraping service, SerpApi. The measure in question was a challenge system standing in front of Google's search results, and the court treated it as a measure that controls access. So a file that asks is below the line, and a challenge you must pass before the page is served is above it. The rungs in between are not yet lit.
Two cautions, because the July ruling is young. It decided a motion to dismiss, so the court took Reddit's allegations as true and nothing has been proved yet. And the challenge that was got around belonged to Google, not to Reddit: if that holds, the person who puts up the wall and the person who owns the work behind it no longer have to be the same.
3Where a careful crawler stops: a worked example
UnGovr, the organisation behind LexLint, runs its own crawler, and its rules show the ladder in practice. The crawler identifies itself as UnGovrBot and signs its requests under a published internet standard, with its public key where any server can fetch it, so a site always knows who is asking. It never goes past a robots.txt refusal.
A bot challenge is different: the crawler may try again through its own tiers, from a plain request up to a full browser. That is the rung the July 2026 ruling put in play. The ruling was about a challenge in front of copyrighted work, put up by a third party, and it was decided at the motion-to-dismiss stage, so a different site, a different work or a different court could read the same rung another way. What one operator does there is a risk it has chosen to carry, not a line the law has drawn, and your project needs its own answer. A CAPTCHA is where it stops. And there are sites it will only ever visit as the declared crawler, because on those sites being recognised matters more than getting the page. The crawler's own page is https://www.ungovr.org/crawler.
The long read. 403 Forbidden: which barriers the law will actually defend carries the depth behind this document, with every citation and its date.
The long read. Disallow: / - what a website can actually forbid carries the depth behind this document, with every citation and its date.
The long read. By continuing, you agree: when terms of use bind a crawler carries the depth behind this document, with every citation and its date.