[{"data":1,"prerenderedAt":206},["ShallowReactive",2],{"learn-lesson-web-scraping-legal-and-ethical-is-web-scraping-legal":3},{"course":4,"lesson":66,"index":179,"outline":180,"prev":204,"next":205},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":34,"syllabus":45,"faq":48,"lessonCount":65},"web-scraping-legal-and-ethical",6,"No legal background assumed","5 lessons, about 55 minutes","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Web Scraping Legal and Ethical: A Free 5-Lesson Course","Is web scraping legal? Public data, terms of service, logins, personal data and GDPR, robots.txt and rate limits, and how to brief your legal team. Written for practitioners, not lawyers. Free, ungated.","is web scraping legal, web scraping gdpr, terms of service scraping, robots txt, public data scraping, web scraping ethics, scraping personal data, scraping compliance","A free course on the legal and ethical side of web scraping","Five written lessons on what you can collect, what changes when you log in, where personal data rules apply, and how to answer your legal team.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course six","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.",{"label":21,"url":22},"Start with lesson one","/learn/web-scraping-legal-and-ethical/is-web-scraping-legal",{"label":24,"url":25},"See what we collect and publish","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33],"Separate the three distinct questions people collapse into \"is scraping legal\"","Tell the difference between public data, terms-bound data and data behind a login, and why that line matters more than any other","Recognise when personal data rules apply to a dataset you thought was about products","Set rate limits and identification that you would be comfortable defending in writing","Produce a one-page brief that gets a useful answer from your legal team instead of a reflexive no",{"title":35,"for_title":36,"for":37,"not_title":41,"not_for":42},"Who this is for","Written for",[38,39,40],"Anyone who has been asked \"are we allowed to do that?\" and does not have a good answer ready","Developers and analysts who want to make a defensible decision rather than an optimistic one","Teams preparing to put a data collection project in front of legal, procurement or a customer's security review","Not written for",[43,44],"Anyone needing an authoritative legal opinion. This is background so that the conversation with a qualified lawyer is a short one.","Readers looking for a jurisdiction-by-jurisdiction reference. The principles here travel; the specifics do not.",{"title":46,"intro":47},"The five lessons","Lesson two contains the distinction that decides most real cases. Lesson five is the deliverable — if you only read one, read that.",{"badge":49,"title":50,"description":51,"items":52},"FAQ","Before you start","The questions that come up in the first meeting, every time.",[53,56,59,62],{"title":54,"description":55},"Is this legal advice?","No, and it cannot be. It is a practitioner's map of the questions that matter, written so that when you do speak to a qualified lawyer in your jurisdiction, you arrive with a specific description rather than \"can we scrape?\". That conversation is much shorter and much cheaper when the facts are already written down.",{"title":57,"description":58},"So is web scraping legal or not?","The question is too coarse to have an answer. Collecting publicly posted prices, respectfully, for market analysis sits in a very different place from harvesting personal profiles from behind a login. Lesson one breaks the question into the three separate ones it actually contains.",{"title":60,"description":61},"Does robots.txt have legal force?","It is a convention rather than a contract, and treating it as either irrelevant or binding both miss the point. Lesson four covers what it is for and why ignoring it is a bad idea regardless of what a court would say about it.",{"title":63,"description":64},"What does Scrapewise itself refuse to do?","We do not collect from behind logins, we do not build personal profile datasets, and we rate-limit by default. Lesson five includes the acceptable-use position we actually operate under, because a vendor who will not tell you where their line is has not thought about it.",5,{"slug":67,"nav_title":68,"title":69,"summary":70,"time":71,"needs_account":72,"seo":73,"blocks":77,"takeaways":170,"next_step":175},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.","10 min",false,{"title":74,"description":75,"keywords":76},"Is Web Scraping Legal? The Three Questions That Matter","Access, copying and use are three separate legal questions that get collapsed into one. How to tell them apart and why most scraping disputes turn on only one of them.","is web scraping legal, web scraping law, scraping legality, data collection compliance, database rights",[78,83,103,109,115,120,132,156,165],{"type":79,"paragraphs":80},"prose",[81,82],"Nobody can answer \"is web scraping legal\" because it is not one question. It is at least three, they are governed by different bodies of law, and a project can be comfortably fine on two of them and in trouble on the third.","Pulling them apart is not a lawyerly evasion — it is the thing that turns an unanswerable worry into a set of specific, checkable facts.",{"type":84,"title":85,"headers":86,"rows":90},"table","The three questions",[87,88,89],"Question","What it is about","What usually decides it",[91,95,99],[92,93,94],"Access","Were you allowed to request the page at all?","Logins, access controls, and whether you circumvented anything",[96,97,98],"Copying","May you store and reproduce what came back?","Copyright in the content, and database rights in the collection",[100,101,102],"Use","May you do the thing you intend with it?","Personal data rules, competition law, contract terms",{"type":79,"title":104,"paragraphs":105},"Access is where the sharpest line sits",[106,107,108],"Requesting a page that any member of the public can load without signing in is, in most jurisdictions, close to the uncontroversial end. You are doing what a browser does, faster.","Everything changes at the login. Once there is an account, there is an agreement you accepted, and there is an access control you are operating inside rather than outside. That shifts the question from \"did you read a public page\" to \"did you comply with the terms you agreed to\", which is a contract question with a much clearer answer and much clearer consequences.","This single distinction — public page versus authenticated session — does more work than any other idea in this course. It is why lesson two spends its whole length on it.",{"type":79,"title":110,"paragraphs":111},"Copying is about the collection more than the item",[112,113,114],"A price is a fact, and facts are not generally protected by copyright. That is why price monitoring sits on firmer ground than people assume.","Two things change that. Creative content — product photography, written descriptions, reviews — is somebody's work and copying it wholesale is a different proposition from recording a number. And in the EU and UK there is a separate database right protecting substantial investment in compiling a collection, which can apply even where no individual item is protected. Taking a substantial part of someone's assembled catalogue is a different act from checking the price of forty products.","The practical consequence is that scope matters. Collecting the fields you need for a specific purpose is a narrower and more defensible act than mirroring a catalogue because you could.",{"type":116,"variant":117,"title":118,"text":119},"callout","note","Purpose changes the answer more than technique does","The same requests, producing the same rows, sit very differently depending on what happens next. Monitoring competitor prices to set your own is ordinary commercial activity that long predates the web. Reproducing a competitor's catalogue as your own storefront is not, and no amount of care in how you fetched the pages makes it so. When you write the project down, lead with the purpose.",{"type":121,"title":122,"intro":123,"items":124},"list","The factors that move a project towards the comfortable end","None is decisive alone. Together they are what a defensible position looks like.",[125,126,127,128,129,130,131],"No login, no access control, nothing circumvented","Facts rather than creative content — prices, availability, specifications","A narrow, stated purpose, and only the fields that purpose needs","A request rate that imposes no meaningful cost on the source","No personal data, or a deliberate decision about it (lesson three)","An identifiable user agent and a contact address","A written record of all of the above, made before you started",{"type":84,"title":133,"intro":134,"headers":135,"rows":138},"Worked example: one project, split into the three questions","A pricing team wants to track 1,200 of its own SKUs across five competitor sites. Written as one question it sounds unanswerable. Split into access, copying and use, every row has a plain answer and the one row that needs a decision becomes obvious.",[87,136,137],"What it actually asks here","Answer for this project",[139,142,145,148,152],[92,140,141],"Are we getting the pages the way an ordinary visitor does, without an account, without a password, without stepping around a block?","Yes. Public category and product pages, no login, standard requests, no bypass of any gate.",[96,143,144],"How much of each page do we keep, and could the stored collection substitute for the source?","Eight fields per product. No descriptions, no images, no reviews. Nobody could shop from our table.",[100,146,147],"What do we do with it, and does that compete with the source's own use of it?","Internal repricing input. Not republished, not resold, not shown to customers.",[149,150,151],"Retention","How long do we keep it, and do we still need the oldest rows?","13 months of daily snapshots, then aggregate to weekly. Nothing beyond that.",[153,154,155],"The row needing a decision","Two of the five sites require an account to see trade prices.","Those two go to the legal brief separately. The other three do not need one.",{"type":121,"title":157,"intro":158,"items":159},"What usually goes wrong","Almost every uncomfortable scraping project got there by one of these, not by a surprise in the law.",[160,161,162,163,164],"Asking \"is scraping legal\" as a single question, getting a shrug, and treating the shrug as permission.","Letting the field list grow quietly. Eight columns is a price feed. Add descriptions, images and reviews and you have built a copy of the catalogue, which is a different conversation.","Starting with the hardest site. The two login-gated competitors drag the other three into a review they never needed.","Assuming the technique decides the answer. The same HTTP request is unremarkable for internal price comparison and awkward for a public mirror of someone's catalogue.","Never writing the three answers down, so when someone asks six months later nobody can reconstruct what was decided or when.",{"type":79,"title":166,"paragraphs":167},"The honest summary",[168,169],"Collecting publicly available factual data, at a considerate rate, for your own analysis, is routine and widespread. Enormous parts of the modern web — search engines, price comparison, academic research, archiving — depend on it being so.","The cases that go wrong cluster tightly, and the clusters are recognisable: data behind a login, personal data at scale, volumes that cost the source real money, and republishing someone's collection as your own. If your project is in none of those clusters, you are in ordinary territory. If it is in one of them, you need advice specific to your jurisdiction before you start, not after.",[171,172,173,174],"Access, copying and use are three separate questions with three different bodies of law behind them.","The public-page versus logged-in line is the sharpest one and does the most work.","Facts are weakly protected; creative content and assembled databases are not. Scope narrowly.","The problem cases cluster: logins, personal data at scale, high cost to the source, republishing a catalogue.",{"text":176,"label":177,"url":178},"Next: the line that matters most — what a terms-of-service page does, and what changes at the login.","Lesson 2: public data, terms of service and logins","/learn/web-scraping-legal-and-ethical/public-data-terms-of-service-and-logins",0,[181,182,188,193,198],{"slug":67,"navTitle":68,"title":69,"summary":70,"time":71,"needsAccount":72},{"slug":183,"navTitle":184,"title":185,"summary":186,"time":187,"needsAccount":72},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.","11 min",{"slug":189,"navTitle":190,"title":191,"summary":192,"time":187,"needsAccount":72},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"slug":194,"navTitle":195,"title":196,"summary":197,"time":187,"needsAccount":72},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.",{"slug":199,"navTitle":200,"title":201,"summary":202,"time":203,"needsAccount":72},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.","12 min",null,{"slug":183,"navTitle":184,"title":185,"summary":186,"time":187,"needsAccount":72},1791047867421]