[{"data":1,"prerenderedAt":239},["ShallowReactive",2],{"learn-lesson-web-scraping-legal-and-ethical-what-to-put-in-front-of-your-legal-team":3},{"course":4,"lesson":66,"index":212,"outline":213,"prev":237,"next":238},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":34,"syllabus":45,"faq":48,"lessonCount":65},"web-scraping-legal-and-ethical",6,"No legal background assumed","5 lessons, about 55 minutes","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Web Scraping Legal and Ethical: A Free 5-Lesson Course","Is web scraping legal? Public data, terms of service, logins, personal data and GDPR, robots.txt and rate limits, and how to brief your legal team. Written for practitioners, not lawyers. Free, ungated.","is web scraping legal, web scraping gdpr, terms of service scraping, robots txt, public data scraping, web scraping ethics, scraping personal data, scraping compliance","A free course on the legal and ethical side of web scraping","Five written lessons on what you can collect, what changes when you log in, where personal data rules apply, and how to answer your legal team.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course six","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.",{"label":21,"url":22},"Start with lesson one","/learn/web-scraping-legal-and-ethical/is-web-scraping-legal",{"label":24,"url":25},"See what we collect and publish","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33],"Separate the three distinct questions people collapse into \"is scraping legal\"","Tell the difference between public data, terms-bound data and data behind a login, and why that line matters more than any other","Recognise when personal data rules apply to a dataset you thought was about products","Set rate limits and identification that you would be comfortable defending in writing","Produce a one-page brief that gets a useful answer from your legal team instead of a reflexive no",{"title":35,"for_title":36,"for":37,"not_title":41,"not_for":42},"Who this is for","Written for",[38,39,40],"Anyone who has been asked \"are we allowed to do that?\" and does not have a good answer ready","Developers and analysts who want to make a defensible decision rather than an optimistic one","Teams preparing to put a data collection project in front of legal, procurement or a customer's security review","Not written for",[43,44],"Anyone needing an authoritative legal opinion. This is background so that the conversation with a qualified lawyer is a short one.","Readers looking for a jurisdiction-by-jurisdiction reference. The principles here travel; the specifics do not.",{"title":46,"intro":47},"The five lessons","Lesson two contains the distinction that decides most real cases. Lesson five is the deliverable — if you only read one, read that.",{"badge":49,"title":50,"description":51,"items":52},"FAQ","Before you start","The questions that come up in the first meeting, every time.",[53,56,59,62],{"title":54,"description":55},"Is this legal advice?","No, and it cannot be. It is a practitioner's map of the questions that matter, written so that when you do speak to a qualified lawyer in your jurisdiction, you arrive with a specific description rather than \"can we scrape?\". That conversation is much shorter and much cheaper when the facts are already written down.",{"title":57,"description":58},"So is web scraping legal or not?","The question is too coarse to have an answer. Collecting publicly posted prices, respectfully, for market analysis sits in a very different place from harvesting personal profiles from behind a login. Lesson one breaks the question into the three separate ones it actually contains.",{"title":60,"description":61},"Does robots.txt have legal force?","It is a convention rather than a contract, and treating it as either irrelevant or binding both miss the point. Lesson four covers what it is for and why ignoring it is a bad idea regardless of what a court would say about it.",{"title":63,"description":64},"What does Scrapewise itself refuse to do?","We do not collect from behind logins, we do not build personal profile datasets, and we rate-limit by default. Lesson five includes the acceptable-use position we actually operate under, because a vendor who will not tell you where their line is has not thought about it.",5,{"slug":67,"nav_title":68,"title":69,"summary":70,"time":71,"needs_account":72,"seo":73,"blocks":77,"takeaways":204,"next_step":209},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.","12 min",false,{"title":74,"description":75,"keywords":76},"How to Get Legal Approval for a Web Scraping Project","A one-page brief template for data collection projects, the three framings that guarantee a reflexive no, and an example acceptable-use position.","web scraping legal approval, data collection compliance brief, legitimate interests assessment, scraping policy, acceptable use data collection",[78,83,112,117,137,143,149,181,192],{"type":79,"paragraphs":80},"prose",[81,82],"Most data collection projects that die in legal review do not die because the answer was no. They die because the question was unanswerable, so the safe response was no.","\"Are we allowed to scrape competitor websites?\" has no good answer. A lawyer hearing it has to imagine the worst version of what you might mean and advise against that. One page of specifics changes the entire conversation, and writing it takes about an hour.",{"type":84,"title":85,"intro":86,"items":87},"steps","The one-page brief","Eight sections, a few sentences each. Written before you start, not after someone asks.",[88,91,94,97,100,103,106,109],{"title":89,"text":90},"Purpose, in one sentence","What business decision this data informs. \"Weekly price position against eleven named competitors on the four hundred SKUs we actively compete on.\" Specific, bounded, obviously ordinary.",{"title":92,"text":93},"Sources, named","The actual domains. Not \"competitor websites\" — the list. Vagueness here reads as evasion even when it is laziness.",{"title":95,"text":96},"Fields, exhaustively","Every column you will store. This is where you demonstrate minimisation, and where a reviewer can see at a glance that there is no personal data in it.",{"title":98,"text":99},"Access method","Public pages, no login, nothing circumvented. State it plainly, because it is the question they most want answered.",{"title":101,"text":102},"Terms review","What each source's terms say about automated access, and the date you read them. A table with one row per source.",{"title":104,"text":105},"Rate and frequency","Requests per minute, concurrency, how often the run happens, and how that compares to the source's own traffic. Numbers, not adjectives.",{"title":107,"text":108},"Personal data position","Either \"none collected, here is the field list\" — much the better answer — or your lawful basis and retention period.",{"title":110,"text":111},"Who sees the output","Internal only, a specific team, or something customer-facing. Republishing changes the analysis completely and is the thing they will ask about if you do not say.",{"type":113,"variant":114,"title":115,"text":116},"callout","note","The date on the terms review is doing real work","It shows a process rather than an assumption, and it gives you a defensible position if the terms change later: you checked, on a date, and recorded what they said. A reviewer reading that knows they are dealing with someone who will also notice when something changes.",{"type":118,"title":119,"headers":120,"rows":124},"table","Three framings that guarantee a no",[121,122,123],"What people say","What it sounds like","Say this instead",[125,129,133],[126,127,128],"\"It's all public data\"","You have not thought about personal data or terms","\"No login, no personal data, here is the field list\"",[130,131,132],"\"Everyone does it\"","You have no position of your own","\"This is standard market research; here is our conduct policy\"",[134,135,136],"\"They can't tell it's us\"","You are relying on not being caught","\"We identify ourselves and publish a contact address\"",{"type":79,"title":138,"paragraphs":139},"Expect conditions, not a verdict",[140,141,142],"A good review rarely produces a clean yes. It produces a yes with conditions, and the conditions are usually reasonable and cheap: drop two fields, halve the frequency, exclude one source whose terms are explicit, add a retention limit, re-check the terms annually.","Treat that as the success case. A conditional approval is a documented, bounded authorisation that protects you and the project, and it is far more valuable than an unconditional shrug that nobody wrote down.","If the answer is a flat no, ask which of the eight sections caused it. Frequently it is one — a single source, or one field — and the project survives without it.",{"type":79,"title":144,"paragraphs":145},"The position we operate under",[146,147,148],"Worth stating, because a vendor who will not tell you where their line is has not thought about it, and you should ask every vendor this question.","We collect from public pages only. No logins, no credential sharing, no circumvention of access controls. We rate-limit by default and back off on 429 and 503. We do not build datasets about individuals — no seller names, no review text, no contact details. We honour removal requests from site operators. And we tell customers plainly when a source cannot be collected rather than filling the gap with a plausible number, because a coverage gap is a known quantity and an invented row is not.","That last one is a data integrity commitment rather than a legal one, but it belongs in the same list. The projects that cause problems downstream are rarely the ones with a documented gap. They are the ones where somebody decided a gap looked bad.",{"type":118,"title":150,"intro":151,"headers":152,"rows":156},"Worked example: the same project, briefed badly and briefed well","Two versions of one request about the same five competitors. The left column is what legal teams usually receive. The right column is the same project described in terms somebody can actually sign off, line for line.",[153,154,155],"What they usually get","What they can answer","Why the difference matters",[157,161,165,169,173,177],[158,159,160],"\"Can we scrape competitor websites?\"","\"We want to read 1,200 public product pages across five named sites, once a day, off-peak, at roughly 1 request per second.\"","The first has no boundary, so the safe answer is no. The second has five facts to check.",[162,163,164],"\"We'd collect pricing data.\"","\"Eight fields: title, GTIN, MPN, price, currency, availability, shipping cost, URL. No descriptions, no images, no reviews, no seller names.\"","A named field list is the single strongest thing in the brief. It proves the collection cannot substitute for the source.",[166,167,168],"\"For analysis.\"","\"Input to internal repricing. Not republished, not resold, not shown to customers or in marketing.\"","Use decides more of the answer than technique does.",[170,171,172],"\"It's all public.\"","\"Three sites need no account. Two show trade prices only behind a login, and we have excluded those two pending your view.\"","Flagging the hard case yourself is what makes the rest credible.",[174,175,176],"(no date)","\"Terms of each site reviewed on 14 March. We re-review every six months and on any redesign.\"","An undated review is an assertion. A dated one is a control.",[178,179,180],"(no exit)","\"We stop within one business day of any request from the site operator, and here is the mailbox that receives it.\"","Reversibility turns a permanent decision into a revocable one.",{"type":182,"title":183,"intro":184,"items":185},"list","What usually goes wrong","A no from legal is usually a response to the brief, not to the project.",[186,187,188,189,190,191],"Asking the abstract question. \"Is scraping legal\" has no answer that helps anyone, and the only safe response to an unbounded question is no.","Asking after the crawler is already running, which turns an approval into an incident review.","Hiding the login-gated sites in the middle of the list instead of naming them as the open question.","Promising a field list and then letting it grow, so the thing that was approved and the thing that runs drift apart within a quarter.","Treating the answer as permanent. Conditions expire, sites get redesigned, terms change, and a review with no date on it stops being evidence of anything.","Leaving no owner. If no named person re-reviews on a schedule, the brief is a document rather than a control.",{"type":182,"title":193,"intro":194,"items":195},"The closing checklist","If you can tick all of these, you are in ordinary territory and you have the paperwork to show it.",[196,197,198,199,200,201,202,203],"Purpose written in one sentence, and it does not sound evasive","Public pages only, no account, nothing circumvented","Field list is the minimum the purpose needs, and contains no personal data","Terms reviewed per source, with dates recorded","Rate and frequency stated as numbers, and justified against how often the data changes","Identifiable user agent with a working contact address","Someone qualified has read the brief and signed off, with any conditions written down","A review date in the calendar — terms change, catalogues change, and so does your purpose",[205,206,207,208],"Projects fail review because the question was unanswerable, not because the answer was no.","Eight sections, one page, written before you start. Named sources, exhaustive field list, numbers for rate and frequency.","\"It's public\", \"everyone does it\" and \"they can't tell it's us\" each guarantee a no. Replace each with a specific.","Expect conditional approval and treat it as the win. Ask every vendor where their own line is.",{"text":210,"label":211,"url":25},"That is the end of the course. If you want to see what collected data actually looks like before committing to anything, every retailer page publishes real measured output — including the ones where we say plainly that nothing readable came back.","See real run output by retailer",4,[214,220,226,231,236],{"slug":215,"navTitle":216,"title":217,"summary":218,"time":219,"needsAccount":72},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.","10 min",{"slug":221,"navTitle":222,"title":223,"summary":224,"time":225,"needsAccount":72},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.","11 min",{"slug":227,"navTitle":228,"title":229,"summary":230,"time":225,"needsAccount":72},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"slug":232,"navTitle":233,"title":234,"summary":235,"time":225,"needsAccount":72},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.",{"slug":67,"navTitle":68,"title":69,"summary":70,"time":71,"needsAccount":72},{"slug":232,"navTitle":233,"title":234,"summary":235,"time":225,"needsAccount":72},null,1791047867487]