[{"data":1,"prerenderedAt":227},["ShallowReactive",2],{"learn-lesson-web-scraping-legal-and-ethical-rate-limits-robots-and-being-a-good-citizen":3},{"course":4,"lesson":66,"index":200,"outline":201,"prev":225,"next":226},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":34,"syllabus":45,"faq":48,"lessonCount":65},"web-scraping-legal-and-ethical",6,"No legal background assumed","5 lessons, about 55 minutes","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"Web Scraping Legal and Ethical: A Free 5-Lesson Course","Is web scraping legal? Public data, terms of service, logins, personal data and GDPR, robots.txt and rate limits, and how to brief your legal team. Written for practitioners, not lawyers. Free, ungated.","is web scraping legal, web scraping gdpr, terms of service scraping, robots txt, public data scraping, web scraping ethics, scraping personal data, scraping compliance","A free course on the legal and ethical side of web scraping","Five written lessons on what you can collect, what changes when you log in, where personal data rules apply, and how to answer your legal team.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course six","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.",{"label":21,"url":22},"Start with lesson one","/learn/web-scraping-legal-and-ethical/is-web-scraping-legal",{"label":24,"url":25},"See what we collect and publish","/custom-scrapers",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33],"Separate the three distinct questions people collapse into \"is scraping legal\"","Tell the difference between public data, terms-bound data and data behind a login, and why that line matters more than any other","Recognise when personal data rules apply to a dataset you thought was about products","Set rate limits and identification that you would be comfortable defending in writing","Produce a one-page brief that gets a useful answer from your legal team instead of a reflexive no",{"title":35,"for_title":36,"for":37,"not_title":41,"not_for":42},"Who this is for","Written for",[38,39,40],"Anyone who has been asked \"are we allowed to do that?\" and does not have a good answer ready","Developers and analysts who want to make a defensible decision rather than an optimistic one","Teams preparing to put a data collection project in front of legal, procurement or a customer's security review","Not written for",[43,44],"Anyone needing an authoritative legal opinion. This is background so that the conversation with a qualified lawyer is a short one.","Readers looking for a jurisdiction-by-jurisdiction reference. The principles here travel; the specifics do not.",{"title":46,"intro":47},"The five lessons","Lesson two contains the distinction that decides most real cases. Lesson five is the deliverable — if you only read one, read that.",{"badge":49,"title":50,"description":51,"items":52},"FAQ","Before you start","The questions that come up in the first meeting, every time.",[53,56,59,62],{"title":54,"description":55},"Is this legal advice?","No, and it cannot be. It is a practitioner's map of the questions that matter, written so that when you do speak to a qualified lawyer in your jurisdiction, you arrive with a specific description rather than \"can we scrape?\". That conversation is much shorter and much cheaper when the facts are already written down.",{"title":57,"description":58},"So is web scraping legal or not?","The question is too coarse to have an answer. Collecting publicly posted prices, respectfully, for market analysis sits in a very different place from harvesting personal profiles from behind a login. Lesson one breaks the question into the three separate ones it actually contains.",{"title":60,"description":61},"Does robots.txt have legal force?","It is a convention rather than a contract, and treating it as either irrelevant or binding both miss the point. Lesson four covers what it is for and why ignoring it is a bad idea regardless of what a court would say about it.",{"title":63,"description":64},"What does Scrapewise itself refuse to do?","We do not collect from behind logins, we do not build personal profile datasets, and we rate-limit by default. Lesson five includes the acceptable-use position we actually operate under, because a vendor who will not tell you where their line is has not thought about it.",5,{"slug":67,"nav_title":68,"title":69,"summary":70,"time":71,"needs_account":72,"seo":73,"blocks":77,"takeaways":191,"next_step":196},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.","11 min",false,{"title":74,"description":75,"keywords":76},"Robots.txt and Scraping Rate Limits: Good Practice","What robots.txt is and is not, how to choose a request rate, why identifying your crawler beats impersonating a browser, and the conduct that keeps you unblocked.","robots txt scraping, scraping rate limit, crawl delay, scraper user agent, web scraping ethics, polite crawling",[78,83,89,94,124,130,136,147,176,186],{"type":79,"paragraphs":80},"prose",[81,82],"The previous three lessons were about what you are permitted to do. This one is about how you do it, which is separate and in day-to-day terms more consequential — because conduct is what determines whether you stay unblocked, and whether a complaint ever reaches anybody's desk.","It is also the part most within your control. You usually cannot change whether a site has a login. You can always change how often you ask.",{"type":79,"title":84,"paragraphs":85},"What robots.txt actually is",[86,87,88],"A text file at the root of a domain, in which the site operator states which paths they would prefer automated clients not to request. It is a convention from the early web, it is honoured voluntarily, and it is not a technical control — nothing stops you ignoring it.","Whether it carries legal weight varies and is genuinely unsettled. The practical case for respecting it does not depend on that. It is the only channel a site has for saying what it wants, it costs almost nothing to honour, and disregarding it is the clearest possible evidence of bad faith if anybody ever looks at your behaviour.","Read it with some judgement, though. A blanket disallow aimed at search engine indexing is a different statement from a specific disallow on a product path. If a file forbids everything and you believe your use is legitimate, the right move is to ask the operator — not to decide unilaterally that the file did not mean you.",{"type":90,"variant":91,"title":92,"text":93},"callout","note","Honour the crawl-delay even though nobody does","Where robots.txt specifies a crawl-delay, it is the site telling you exactly what rate it is comfortable with. Taking them at their word costs you a slower run and buys you a position that is very hard to criticise. It is the cheapest good-faith evidence available and it is routinely ignored.",{"type":95,"title":96,"intro":97,"headers":98,"rows":103},"table","Choosing a rate","Rules of thumb, not standards. When in doubt, be slower — nobody has ever been blocked for being too polite.",[99,100,101,102],"Site type","Reasonable concurrency","Reasonable gap","Why",[104,109,114,119],[105,106,107,108],"Large retailer or marketplace","2 to 4","1 to 2 seconds","They serve far more than this per second already",[110,111,112,113],"Mid-sized e-commerce","1 to 2","2 to 5 seconds","Your traffic is visible in their analytics",[115,116,117,118],"Small or independent site","1","5 to 10 seconds","You could be a noticeable share of their load",[120,121,122,123],"Anything that has slowed down","Back off immediately","Exponential","A 429 or a rising latency is a request to stop",{"type":79,"title":125,"paragraphs":126},"Frequency is a separate question from rate, and it is the bigger one",[127,128,129],"Rate is how fast you go through a run. Frequency is how often you do the run, and it is where most unnecessary load comes from.","The honest test: how often does the underlying value actually change? A price that moves twice a week does not need hourly collection, and most schedules are set by what felt responsive at the time rather than by any measurement. Halving your frequency halves your load on the source, halves your cost, and in most cases changes nothing anybody downstream would notice.","Better still, make the frequency match the volatility. Fast-moving categories daily, stable ones weekly. This is both more considerate and cheaper, which is a rare combination.",{"type":79,"title":131,"paragraphs":132},"Identify yourself",[133,134,135],"The default instinct is to blend in — copy a browser's user agent and hope nobody looks. It is understandable and it is usually the wrong call.","A user agent naming your organisation with a contact URL means that when somebody does look at their logs, they find a named, contactable, well-behaved client. The realistic outcomes are an allowlist, an email asking you to slow down, or nothing at all. The outcomes available to an anonymous client failing a fingerprint check are a ban and, if it ever escalates, a much worse story.","There is a real tension here, because some protection layers are more hostile to a declared crawler than to a convincing browser impersonation. Our position is that the long game favours being identifiable: impersonation is an arms race you lose eventually, and a relationship is durable in a way that a working header set is not.",{"type":137,"title":138,"items":139},"list","Conduct that keeps you welcome",[140,141,142,143,144,145,146],"Back off on 429 and 503 rather than retrying immediately — those are explicit requests to slow down","Run heavy jobs outside the source's peak hours","Cache, and never request a page twice when once would do","Request only what you need — no images, no assets, no pages outside the scope","Publish a contact address and answer it","Stop promptly if asked, and talk to them before resuming","Review the schedule periodically. Jobs set up two years ago are usually running far more often than anyone currently needs.",{"type":95,"title":148,"intro":149,"headers":150,"rows":155},"Worked example: the same 38,600 pages at four different rates","One competitor, 38,600 product URLs, run once a day. The rate you pick decides both how long the run takes and how visible you are in their logs. Their own traffic is the number that matters: a mid-size retailer serving roughly 40,000 page views a day is handling about 0.5 requests per second on average.",[151,152,153,154],"Rate","Run duration","Share of their average traffic","How it reads in their logs",[156,161,166,171],[157,158,159,160],"10 req/s","64 minutes","About 20x their own average rate","A spike. Someone will look, and they should.",[162,163,164,165],"2 req/s","5 hours 22 minutes","About 4x","Still the loudest client they have that hour.",[167,168,169,170],"0.5 req/s","21 hours 26 minutes","About 1x","Indistinguishable from a busy customer, but it no longer fits in a day.",[172,173,174,175],"1 req/s, split across 2 nightly windows","10 hours 43 minutes, off-peak","About 2x, during their quietest hours","Invisible in practice. This is the one to pick.",{"type":137,"title":177,"intro":178,"items":179},"What usually goes wrong","Most of the damage comes from the schedule rather than the rate.",[180,181,182,183,184,185],"Tuning the rate carefully and then running six scrapers in parallel against the same host, so the real figure is six times the one on the dashboard.","Running hourly because the scheduler made it easy, when prices on that site change twice a week. Frequency multiplies everything.","Retrying failures immediately, so the moment a site is struggling is exactly the moment your traffic triples.","Running at the top of the hour like everyone else's cron, which stacks your load onto theirs.","Using a generic browser user agent with no contact address, which turns a solvable conversation into a block. Being identifiable is what lets someone email you instead of banning you.","Scraping the same page five times to get five fields, because the extraction was written field by field rather than page by page.",{"type":79,"title":187,"paragraphs":188},"The underlying principle",[189,190],"Be the kind of client a site operator would not bother blocking, because blocking you would be more effort than tolerating you.","That is not only an ethical stance, it is the most reliable operational strategy available. The scrapers that get blocked are the ones that are expensive, anonymous and relentless. The ones that run for years are polite, identifiable and ask for exactly what they need.",[192,193,194,195],"robots.txt is a convention, not a control, and honouring it costs almost nothing while disregarding it is clear evidence of bad faith.","Be slower than you think you need. Nobody has ever been blocked for politeness.","Frequency matters more than rate — match collection frequency to how often the value actually changes.","Identify yourself with a contact address. Impersonation is an arms race you lose; a relationship is durable.",{"text":197,"label":198,"url":199},"Finally: turning all of this into the one page that gets you an answer instead of a reflexive no.","Lesson 5: what to put in front of your legal team","/learn/web-scraping-legal-and-ethical/what-to-put-in-front-of-your-legal-team",3,[202,208,213,218,219],{"slug":203,"navTitle":204,"title":205,"summary":206,"time":207,"needsAccount":72},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.","10 min",{"slug":209,"navTitle":210,"title":211,"summary":212,"time":71,"needsAccount":72},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.",{"slug":214,"navTitle":215,"title":216,"summary":217,"time":71,"needsAccount":72},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"slug":67,"navTitle":68,"title":69,"summary":70,"time":71,"needsAccount":72},{"slug":220,"navTitle":221,"title":222,"summary":223,"time":224,"needsAccount":72},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.","12 min",{"slug":214,"navTitle":215,"title":216,"summary":217,"time":71,"needsAccount":72},{"slug":220,"navTitle":221,"title":222,"summary":223,"time":224,"needsAccount":72},1791047867474]