[{"data":1,"prerenderedAt":214},["ShallowReactive",2],{"learn-lesson-web-scraping-legal-and-ethical-public-data-terms-of-service-and-logins":3},{"course":4,"lesson":66,"index":187,"outline":188,"prev":212,"next":213},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":34,"syllabus":45,"faq":48,"lessonCount":65},"web-scraping-legal-and-ethical",6,"No legal background assumed","5 lessons, about 55 minutes","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Web Scraping Legal and Ethical: A Free 5-Lesson Course","Is web scraping legal? Public data, terms of service, logins, personal data and GDPR, robots.txt and rate limits, and how to brief your legal team. Written for practitioners, not lawyers. Free, ungated.","is web scraping legal, web scraping gdpr, terms of service scraping, robots txt, public data scraping, web scraping ethics, scraping personal data, scraping compliance","A free course on the legal and ethical side of web scraping","Five written lessons on what you can collect, what changes when you log in, where personal data rules apply, and how to answer your legal team.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course six","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.",{"label":21,"url":22},"Start with lesson one","/learn/web-scraping-legal-and-ethical/is-web-scraping-legal",{"label":24,"url":25},"See what we collect and publish","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33],"Separate the three distinct questions people collapse into \"is scraping legal\"","Tell the difference between public data, terms-bound data and data behind a login, and why that line matters more than any other","Recognise when personal data rules apply to a dataset you thought was about products","Set rate limits and identification that you would be comfortable defending in writing","Produce a one-page brief that gets a useful answer from your legal team instead of a reflexive no",{"title":35,"for_title":36,"for":37,"not_title":41,"not_for":42},"Who this is for","Written for",[38,39,40],"Anyone who has been asked \"are we allowed to do that?\" and does not have a good answer ready","Developers and analysts who want to make a defensible decision rather than an optimistic one","Teams preparing to put a data collection project in front of legal, procurement or a customer's security review","Not written for",[43,44],"Anyone needing an authoritative legal opinion. This is background so that the conversation with a qualified lawyer is a short one.","Readers looking for a jurisdiction-by-jurisdiction reference. The principles here travel; the specifics do not.",{"title":46,"intro":47},"The five lessons","Lesson two contains the distinction that decides most real cases. Lesson five is the deliverable — if you only read one, read that.",{"badge":49,"title":50,"description":51,"items":52},"FAQ","Before you start","The questions that come up in the first meeting, every time.",[53,56,59,62],{"title":54,"description":55},"Is this legal advice?","No, and it cannot be. It is a practitioner's map of the questions that matter, written so that when you do speak to a qualified lawyer in your jurisdiction, you arrive with a specific description rather than \"can we scrape?\". That conversation is much shorter and much cheaper when the facts are already written down.",{"title":57,"description":58},"So is web scraping legal or not?","The question is too coarse to have an answer. Collecting publicly posted prices, respectfully, for market analysis sits in a very different place from harvesting personal profiles from behind a login. Lesson one breaks the question into the three separate ones it actually contains.",{"title":60,"description":61},"Does robots.txt have legal force?","It is a convention rather than a contract, and treating it as either irrelevant or binding both miss the point. Lesson four covers what it is for and why ignoring it is a bad idea regardless of what a court would say about it.",{"title":63,"description":64},"What does Scrapewise itself refuse to do?","We do not collect from behind logins, we do not build personal profile datasets, and we rate-limit by default. Lesson five includes the acceptable-use position we actually operate under, because a vendor who will not tell you where their line is has not thought about it.",5,{"slug":67,"nav_title":68,"title":69,"summary":70,"time":71,"needs_account":72,"seo":73,"blocks":77,"takeaways":178,"next_step":183},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.","11 min",false,{"title":74,"description":75,"keywords":76},"Terms of Service and Web Scraping: Browsewrap vs Clickwrap","Why terms you never clicked are weaker than terms you accepted, what changes the moment you create an account, and how that applies to mobile app APIs and consent walls.","terms of service scraping, browsewrap clickwrap, scraping behind login, public data scraping, mobile app api scraping, scraping account terms",[78,82,102,108,113,119,129,163,172],{"type":79,"paragraphs":80},"prose",[81],"Nearly every website has a terms page, nearly all of them say something about automated access, and almost nobody has read the one for the site they are about to collect from. The useful question is not what it says — it is whether you ever agreed to it.",{"type":83,"title":84,"headers":85,"rows":89},"table","Three levels of agreement",[86,87,88],"Situation","How strong is the agreement","Practical consequence",[90,94,98],[91,92,93],"Terms linked in a footer, never clicked","Weakest — often called browsewrap","Contested. Not nothing, but not a clear contract either.",[95,96,97],"A consent or cookie banner you dismissed","Depends entirely on wording and prominence","Grey. Worth reading what you clicked.",[99,100,101],"An account you created, accepting terms","Strongest — clickwrap, a real agreement","You are bound. If it forbids automated access, that is the answer.",{"type":79,"title":103,"paragraphs":104},"Why the distinction is not a loophole",[105,106,107],"It would be convenient to read the first row as \"footer terms don't count\". That is not the point and it is not safe.","The point is that the three rows carry genuinely different weight, and a decision that treats them as identical is making the wrong trade in both directions — either refusing a project that is perfectly ordinary, or walking into a clear contractual breach because \"nobody reads terms\".","A practical middle position: read the automated access clause before you start, record what it said and when you read it, and treat an explicit prohibition as a reason to stop and ask rather than a formality to route around. The record is what turns a judgement call into a documented decision, and documented decisions are what survive a review.",{"type":109,"variant":110,"title":111,"text":112},"callout","warning","Creating an account is a decision, not a convenience","The moment you sign up to see prices more easily, you have accepted terms, you are inside an access control, and your activity is attributable to an identity. If those terms prohibit automated access — and most do — then automating from that account is a breach regardless of how public the underlying pages look. The pragmatic rule we operate under: if the data requires a login, we do not collect it.",{"type":79,"title":114,"paragraphs":115},"Mobile app APIs are not a shortcut around this",[116,117,118],"A recurring discovery is that a site which is hostile to scraping has a mobile app talking to a clean, fast, unprotected JSON endpoint. It is genuinely tempting — better data, less work, no HTML parsing.","It also usually sits worse than the website, not better. The app has its own terms, which you accepted on installation. The endpoint is frequently authenticated, which puts you inside an access control. And obtaining the credentials often involves inspecting traffic in a way that is itself covered by those terms.","There are cases where an app API is openly documented and unauthenticated, and those are fine. The failure mode is assuming that because something is technically reachable, it is in the same category as a public web page. It usually is not, and \"it was easier\" is not a position anybody wants to defend.",{"type":120,"title":121,"intro":122,"items":123},"list","Questions to answer before the first request","Five minutes each, and together they are most of the due diligence anyone will ask you for.",[124,125,126,127,128],"Is this page reachable with no account, no cookie beyond a session, and nothing circumvented?","What does the terms page say about automated access, and on what date did I read it?","Is there a published API or feed that covers the same data? Using it is better in every respect.","Would a reasonable person at that company, seeing exactly what I am doing, consider it a problem?","Can I write down the purpose in one sentence that does not sound evasive?",{"type":83,"title":130,"intro":131,"headers":132,"rows":137},"Worked example: five sites, sorted by what you actually agreed to","The same crawl across five competitors lands in three different places. Sorting the list this way takes about ten minutes and decides which sites need a conversation before the first request.",[133,134,135,136],"Site","How you reach the data","What you agreed to","Where that leaves it",[138,143,148,153,158],[139,140,141,142],"A","Public category pages, no account","Nothing. You never clicked anything.","Proceed. Conduct rules still apply.",[144,145,146,147],"B","Public pages, terms linked in the footer","Nothing you assented to.","Proceed, but read the terms so you know what you are choosing to ignore and can say so out loud.",[149,150,151,152],"C","Public pages behind a cookie banner you must dismiss","Still nothing contractual about data use in most framings, but note it.","Proceed. Record that the banner exists.",[154,155,156,157],"D","Trade prices visible only after creating an account","Everything in the terms you ticked at signup.","Stop. This is a contract question, not a scraping question.",[159,160,161,162],"E","A documented public API with a key","The API terms, plus a rate limit you can actually read.","Use the API. It is the cheapest and clearest of the five.",{"type":120,"title":164,"intro":165,"items":166},"What usually goes wrong","The login line is the one people cross without noticing.",[167,168,169,170,171],"Someone on the team already has an account from a trade show or a test order, so the crawler quietly uses it and nobody records that the project changed category.","Treating \"the data is public once you are logged in\" as the same as public. The account is the thing that changed, not the data.","Reusing a personal account rather than a company one, which moves an individual's name onto a contract they did not read in this context.","Finding the mobile app's unauthenticated endpoint and treating it as a loophole. It is the same site, the same operator, and usually the same terms.","Never re-reading the terms. The sort above is accurate on the day you do it and silently rots after that, which is why the brief in the last lesson carries a date.",{"type":79,"title":173,"paragraphs":174},"The last question is the most useful one",[175,176,177],"The reasonable-person test is not a legal standard, but it is an excellent early-warning system, and it catches things the formal checks miss.","A competitor noticing that you track their public prices will shrug — they almost certainly track yours, and the practice predates e-commerce by decades. A company discovering that you have reconstructed their customer list, or that your collection is measurably slowing their site, will not.","If the honest answer to \"how would they feel about this\" is \"they would be angry, and they would have a point\", that is a signal worth acting on before anybody else gets involved.",[179,180,181,182],"Footer terms, dismissed banners and accepted account terms carry very different weight — do not flatten them.","Read the automated access clause, record what it said and when. The record is the deliverable.","An account turns a public-page question into a contract question. If data needs a login, the safe answer is not to collect it.","Mobile app APIs usually sit worse than the website, not better.",{"text":184,"label":185,"url":186},"Next: the rules that apply even when everything above is settled, because the data turned out to be about people.","Lesson 3: personal data and GDPR","/learn/web-scraping-legal-and-ethical/personal-data-and-gdpr",1,[189,195,196,201,206],{"slug":190,"navTitle":191,"title":192,"summary":193,"time":194,"needsAccount":72},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.","10 min",{"slug":67,"navTitle":68,"title":69,"summary":70,"time":71,"needsAccount":72},{"slug":197,"navTitle":198,"title":199,"summary":200,"time":71,"needsAccount":72},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"slug":202,"navTitle":203,"title":204,"summary":205,"time":71,"needsAccount":72},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.",{"slug":207,"navTitle":208,"title":209,"summary":210,"time":211,"needsAccount":72},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.","12 min",{"slug":190,"navTitle":191,"title":192,"summary":193,"time":194,"needsAccount":72},{"slug":197,"navTitle":198,"title":199,"summary":200,"time":71,"needsAccount":72},1791047867440]