What do I need to worry about? · twelve questions

About this documentUpdated ShowHide

Sean McDermott, Co-Founder and CEO, UnGovr

Written by Sean McDermott (with AI assistance) using the LexLint law library, which supplied every legal instrument, status and date on these pages, and the handbook and insight documents on lexlint.io that carry the depth behind each one.

Every law named here links to its summary page on lexlint.io, translated to English (if needed) and restructured to a standard format for human and code use. Every case links to the court's or the regulator's own record where one could be reached.

© 2026 UnGovr, publishing as LexLint. The text and the figures are licensed under Creative Commons Attribution-ShareAlike 4.0: share and adapt them, including commercially, with credit to LexLint (UnGovr) and under the same licence. Please contact LexLint at hello@ungovr.org to discuss other terms. Logos and wordmarks belong to their owners.

Corpus figures as of .

Legal information, not legal advice. This document describes the law as written and dated; it does not apply it to any system. The notice at the foot says what that means.

Answer twelve questions about what you want to scrape, from where, and what you will do with it. The result is the short list of things to be concerned about, in the order to deal with them, and a profile you can hand to a developer or a lawyer.

1 Where is your organisation based?

The country, and the state or province where the law has one.

2 Where will the scraper actually run?
3 Which sites, and where is the company behind each one?

Add each site. If you do not know where its company is, say so: the second document shows how to find out.

4 Can anyone see the pages you want without logging in?
5 To get the pages, will you create an account, tick a box, or otherwise agree to the site's terms?
6 Does the site say no anywhere you have looked: a robots.txt file, a line in the terms, a “no bots” or “no AI training” notice, a letter?
7 Is there something in the way: a “verify you are human” box, a CAPTCHA, a block after a few pages, a login wall?
8 Does what you want include information about people?

Names, emails, phone numbers, photos, faces, posts, profiles.

9 What will you do with it?
10 Is an AI model or agent doing the collecting or the reading? Whose, and where is that provider based?
11 Will the data leave the country it was collected in?
12 Has anyone from the site told you to stop, or blocked you before?

1What to be concerned about, in order

With the questions above answered, these are the concerns, from the cheapest to deal with to the one that needs a lawyer's view. Every one of them links the document that explains it.

  1. terms of use

    Read the site's terms page before you fetch anything. If it says no to automated access, or to the use you plan, that is the cheapest problem to find and the most expensive to ignore. Read the document.

  2. robots.txt and the other crawl signals

    Read the site's robots.txt and look for a reservation: a “no AI training” line, a TDM header, a licence. In some places it binds you; in most it is evidence against you. Read the document.

  3. terms of use

    The pages are behind a login. You will be inside a contract, and the public-page rule that helps you under computer-misuse law no longer does. Read the document.

  4. where the site is

    You could not say where one or more of the target sites is. Work it out with the evidence ladder, and until you can, assume the stricter of the places it could be and write down that you could not tell. Read the document.

  5. the people in the data

    The pages have people on them. You need a lawful basis, a notice people can find, and a way for them to say no. “It was public” is not the basis. Read the document.

  6. computer misuse and unauthorised access

    Something is in the way. Getting past a technical barrier is where computer-misuse law starts, and in July 2026 a US court treated a challenge in front of a page as a lock under anti-circumvention law. Read the document.

  7. computer misuse and unauthorised access

    You have been told to stop, or blocked. Continuing after a letter or a block breaks the site's terms everywhere, and in many places it turns the next request into unauthorised access; on a public page in the US, the courts are not agreed. Read the document.

  8. copyright and text-and-data mining

    You will publish, sell or train on what you collect. Copying is the act copyright reaches. Whether an exception covers you depends on the place and the use, and in the EU a reservation switches the commercial exception off. Read the document.

  9. database rights

    In the EU and a few other places a compiled database is protected on its own, whatever the copyright in its contents. Taking a substantial part of one is a claim by itself. Read the document.

  10. unfair competition

    You are building something that competes with the site using its own compiled work. That is the fact pattern behind hot-news and unfair-competition claims. Read the document.

  11. AI law

    A model is in the loop. Its provider's place adds that place's law, and a model trained on what you collect brings the AI Act's duty to honour reservations. Read the document.

  12. the people in the data

    The data will leave the country it was collected in. Transfer rules apply on the way out. Read the document.

  13. who runs it for you

    A vendor scrapes for you. Their place and their conduct are yours to answer for in the places you are. Put every question above to them, in writing. Read the document.

2The places

With the script on, each place your answers name appears here, linked to its law page where LexLint has written one. Every place is on the law index.

3The profile, for a developer or a lawyer

With the script on, this section writes a lexlint.yml from your answers and a block to paste into a coding assistant. These are the two lines that install the LexLint plugin.

claude plugin marketplace add ungovr/lexlint && claude plugin install lexlint@lexlint

Legal information, not legal advice. This result lists the obvious problems for the answers you gave. It does not find all of them, and it does not clear a project; that is a question for a lawyer.