The legal and ethical side, without the hand-waving· Lesson 2 of 5

Public data, terms of service, and what changes at the login

Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.

  • 11 min read
  • No account needed

Nearly every website has a terms page, nearly all of them say something about automated access, and almost nobody has read the one for the site they are about to collect from. The useful question is not what it says — it is whether you ever agreed to it.

Three levels of agreement

SituationHow strong is the agreementPractical consequence
Terms linked in a footer, never clickedWeakest — often called browsewrapContested. Not nothing, but not a clear contract either.
A consent or cookie banner you dismissedDepends entirely on wording and prominenceGrey. Worth reading what you clicked.
An account you created, accepting termsStrongest — clickwrap, a real agreementYou are bound. If it forbids automated access, that is the answer.

Why the distinction is not a loophole

It would be convenient to read the first row as "footer terms don't count". That is not the point and it is not safe.

The point is that the three rows carry genuinely different weight, and a decision that treats them as identical is making the wrong trade in both directions — either refusing a project that is perfectly ordinary, or walking into a clear contractual breach because "nobody reads terms".

A practical middle position: read the automated access clause before you start, record what it said and when you read it, and treat an explicit prohibition as a reason to stop and ask rather than a formality to route around. The record is what turns a judgement call into a documented decision, and documented decisions are what survive a review.

Mobile app APIs are not a shortcut around this

A recurring discovery is that a site which is hostile to scraping has a mobile app talking to a clean, fast, unprotected JSON endpoint. It is genuinely tempting — better data, less work, no HTML parsing.

It also usually sits worse than the website, not better. The app has its own terms, which you accepted on installation. The endpoint is frequently authenticated, which puts you inside an access control. And obtaining the credentials often involves inspecting traffic in a way that is itself covered by those terms.

There are cases where an app API is openly documented and unauthenticated, and those are fine. The failure mode is assuming that because something is technically reachable, it is in the same category as a public web page. It usually is not, and "it was easier" is not a position anybody wants to defend.

Questions to answer before the first request

Five minutes each, and together they are most of the due diligence anyone will ask you for.

  • Is this page reachable with no account, no cookie beyond a session, and nothing circumvented?
  • What does the terms page say about automated access, and on what date did I read it?
  • Is there a published API or feed that covers the same data? Using it is better in every respect.
  • Would a reasonable person at that company, seeing exactly what I am doing, consider it a problem?
  • Can I write down the purpose in one sentence that does not sound evasive?

Worked example: five sites, sorted by what you actually agreed to

The same crawl across five competitors lands in three different places. Sorting the list this way takes about ten minutes and decides which sites need a conversation before the first request.

SiteHow you reach the dataWhat you agreed toWhere that leaves it
APublic category pages, no accountNothing. You never clicked anything.Proceed. Conduct rules still apply.
BPublic pages, terms linked in the footerNothing you assented to.Proceed, but read the terms so you know what you are choosing to ignore and can say so out loud.
CPublic pages behind a cookie banner you must dismissStill nothing contractual about data use in most framings, but note it.Proceed. Record that the banner exists.
DTrade prices visible only after creating an accountEverything in the terms you ticked at signup.Stop. This is a contract question, not a scraping question.
EA documented public API with a keyThe API terms, plus a rate limit you can actually read.Use the API. It is the cheapest and clearest of the five.

What usually goes wrong

The login line is the one people cross without noticing.

  • Someone on the team already has an account from a trade show or a test order, so the crawler quietly uses it and nobody records that the project changed category.
  • Treating "the data is public once you are logged in" as the same as public. The account is the thing that changed, not the data.
  • Reusing a personal account rather than a company one, which moves an individual's name onto a contract they did not read in this context.
  • Finding the mobile app's unauthenticated endpoint and treating it as a loophole. It is the same site, the same operator, and usually the same terms.
  • Never re-reading the terms. The sort above is accurate on the day you do it and silently rots after that, which is why the brief in the last lesson carries a date.

The last question is the most useful one

The reasonable-person test is not a legal standard, but it is an excellent early-warning system, and it catches things the formal checks miss.

A competitor noticing that you track their public prices will shrug — they almost certainly track yours, and the practice predates e-commerce by decades. A company discovering that you have reconstructed their customer list, or that your collection is measurably slowing their site, will not.

If the honest answer to "how would they feel about this" is "they would be angry, and they would have a point", that is a signal worth acting on before anybody else gets involved.

Next: the rules that apply even when everything above is settled, because the data turned out to be about people.

Lesson 3: personal data and GDPR

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.