The legal and ethical side, without the hand-waving· Lesson 1 of 5

Three questions hiding inside one

"Is scraping legal" bundles access, copying and use into a single question. Separating them is most of the work.

  • 10 min read
  • No account needed

Nobody can answer "is web scraping legal" because it is not one question. It is at least three, they are governed by different bodies of law, and a project can be comfortably fine on two of them and in trouble on the third.

Pulling them apart is not a lawyerly evasion — it is the thing that turns an unanswerable worry into a set of specific, checkable facts.

The three questions

QuestionWhat it is aboutWhat usually decides it
AccessWere you allowed to request the page at all?Logins, access controls, and whether you circumvented anything
CopyingMay you store and reproduce what came back?Copyright in the content, and database rights in the collection
UseMay you do the thing you intend with it?Personal data rules, competition law, contract terms

Access is where the sharpest line sits

Requesting a page that any member of the public can load without signing in is, in most jurisdictions, close to the uncontroversial end. You are doing what a browser does, faster.

Everything changes at the login. Once there is an account, there is an agreement you accepted, and there is an access control you are operating inside rather than outside. That shifts the question from "did you read a public page" to "did you comply with the terms you agreed to", which is a contract question with a much clearer answer and much clearer consequences.

This single distinction — public page versus authenticated session — does more work than any other idea in this course. It is why lesson two spends its whole length on it.

Copying is about the collection more than the item

A price is a fact, and facts are not generally protected by copyright. That is why price monitoring sits on firmer ground than people assume.

Two things change that. Creative content — product photography, written descriptions, reviews — is somebody's work and copying it wholesale is a different proposition from recording a number. And in the EU and UK there is a separate database right protecting substantial investment in compiling a collection, which can apply even where no individual item is protected. Taking a substantial part of someone's assembled catalogue is a different act from checking the price of forty products.

The practical consequence is that scope matters. Collecting the fields you need for a specific purpose is a narrower and more defensible act than mirroring a catalogue because you could.

The factors that move a project towards the comfortable end

None is decisive alone. Together they are what a defensible position looks like.

  • No login, no access control, nothing circumvented
  • Facts rather than creative content — prices, availability, specifications
  • A narrow, stated purpose, and only the fields that purpose needs
  • A request rate that imposes no meaningful cost on the source
  • No personal data, or a deliberate decision about it (lesson three)
  • An identifiable user agent and a contact address
  • A written record of all of the above, made before you started

Worked example: one project, split into the three questions

A pricing team wants to track 1,200 of its own SKUs across five competitor sites. Written as one question it sounds unanswerable. Split into access, copying and use, every row has a plain answer and the one row that needs a decision becomes obvious.

QuestionWhat it actually asks hereAnswer for this project
AccessAre we getting the pages the way an ordinary visitor does, without an account, without a password, without stepping around a block?Yes. Public category and product pages, no login, standard requests, no bypass of any gate.
CopyingHow much of each page do we keep, and could the stored collection substitute for the source?Eight fields per product. No descriptions, no images, no reviews. Nobody could shop from our table.
UseWhat do we do with it, and does that compete with the source's own use of it?Internal repricing input. Not republished, not resold, not shown to customers.
RetentionHow long do we keep it, and do we still need the oldest rows?13 months of daily snapshots, then aggregate to weekly. Nothing beyond that.
The row needing a decisionTwo of the five sites require an account to see trade prices.Those two go to the legal brief separately. The other three do not need one.

What usually goes wrong

Almost every uncomfortable scraping project got there by one of these, not by a surprise in the law.

  • Asking "is scraping legal" as a single question, getting a shrug, and treating the shrug as permission.
  • Letting the field list grow quietly. Eight columns is a price feed. Add descriptions, images and reviews and you have built a copy of the catalogue, which is a different conversation.
  • Starting with the hardest site. The two login-gated competitors drag the other three into a review they never needed.
  • Assuming the technique decides the answer. The same HTTP request is unremarkable for internal price comparison and awkward for a public mirror of someone's catalogue.
  • Never writing the three answers down, so when someone asks six months later nobody can reconstruct what was decided or when.

The honest summary

Collecting publicly available factual data, at a considerate rate, for your own analysis, is routine and widespread. Enormous parts of the modern web — search engines, price comparison, academic research, archiving — depend on it being so.

The cases that go wrong cluster tightly, and the clusters are recognisable: data behind a login, personal data at scale, volumes that cost the source real money, and republishing someone's collection as your own. If your project is in none of those clusters, you are in ordinary territory. If it is in one of them, you need advice specific to your jurisdiction before you start, not after.

Next: the line that matters most — what a terms-of-service page does, and what changes at the login.

Lesson 2: public data, terms of service and logins

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.