[{"data":1,"prerenderedAt":896},["ShallowReactive",2],{"learn-course-web-scraping-legal-and-ethical":3,"learn-courses":680},{"slug":4,"order":5,"level":6,"time":7,"card_text":8,"seo":9,"hero":15,"outcomes":25,"who":33,"syllabus":44,"faq":47,"lessons":64},"web-scraping-legal-and-ethical",6,"No legal background assumed","5 lessons, about 55 minutes","The question that stops projects: are we allowed to do this? Public data versus terms of service, what changes the moment you log in, where personal data rules bite, what good conduct actually looks like, and how to write the one page your legal team needs.",{"title":10,"description":11,"keywords":12,"og_title":13,"og_description":14},"Web Scraping Legal and Ethical: A Free 5-Lesson Course","Is web scraping legal? Public data, terms of service, logins, personal data and GDPR, robots.txt and rate limits, and how to brief your legal team. Written for practitioners, not lawyers. Free, ungated.","is web scraping legal, web scraping gdpr, terms of service scraping, robots txt, public data scraping, web scraping ethics, scraping personal data, scraping compliance","A free course on the legal and ethical side of web scraping","Five written lessons on what you can collect, what changes when you log in, where personal data rules apply, and how to answer your legal team.",{"badge":16,"title":17,"subtitle":18,"cta_primary":19,"cta_secondary":22},"Course six","The legal and ethical side, without the hand-waving","Most writing on this subject is either a confident \"it's public data, you're fine\" or a lawyer's refusal to say anything useful. Neither helps you decide whether to start. This course sets out the distinctions that actually matter — public versus logged-in, factual versus personal, considerate versus costly — so you can make a defensible call and write it down. It is written by practitioners and it is not legal advice.",{"label":20,"url":21},"Start with lesson one","/learn/web-scraping-legal-and-ethical/is-web-scraping-legal",{"label":23,"url":24},"See what we collect and publish","/custom-scrapers",{"title":26,"items":27},"What you will be able to do",[28,29,30,31,32],"Separate the three distinct questions people collapse into \"is scraping legal\"","Tell the difference between public data, terms-bound data and data behind a login, and why that line matters more than any other","Recognise when personal data rules apply to a dataset you thought was about products","Set rate limits and identification that you would be comfortable defending in writing","Produce a one-page brief that gets a useful answer from your legal team instead of a reflexive no",{"title":34,"for_title":35,"for":36,"not_title":40,"not_for":41},"Who this is for","Written for",[37,38,39],"Anyone who has been asked \"are we allowed to do that?\" and does not have a good answer ready","Developers and analysts who want to make a defensible decision rather than an optimistic one","Teams preparing to put a data collection project in front of legal, procurement or a customer's security review","Not written for",[42,43],"Anyone needing an authoritative legal opinion. This is background so that the conversation with a qualified lawyer is a short one.","Readers looking for a jurisdiction-by-jurisdiction reference. The principles here travel; the specifics do not.",{"title":45,"intro":46},"The five lessons","Lesson two contains the distinction that decides most real cases. Lesson five is the deliverable — if you only read one, read that.",{"badge":48,"title":49,"description":50,"items":51},"FAQ","Before you start","The questions that come up in the first meeting, every time.",[52,55,58,61],{"title":53,"description":54},"Is this legal advice?","No, and it cannot be. It is a practitioner's map of the questions that matter, written so that when you do speak to a qualified lawyer in your jurisdiction, you arrive with a specific description rather than \"can we scrape?\". That conversation is much shorter and much cheaper when the facts are already written down.",{"title":56,"description":57},"So is web scraping legal or not?","The question is too coarse to have an answer. Collecting publicly posted prices, respectfully, for market analysis sits in a very different place from harvesting personal profiles from behind a login. Lesson one breaks the question into the three separate ones it actually contains.",{"title":59,"description":60},"Does robots.txt have legal force?","It is a convention rather than a contract, and treating it as either irrelevant or binding both miss the point. Lesson four covers what it is for and why ignoring it is a bad idea regardless of what a court would say about it.",{"title":62,"description":63},"What does Scrapewise itself refuse to do?","We do not collect from behind logins, we do not build personal profile datasets, and we rate-limit by default. Lesson five includes the acceptable-use position we actually operate under, because a vendor who will not tell you where their line is has not thought about it.",[65,178,293,415,541],{"slug":66,"nav_title":67,"title":68,"summary":69,"time":70,"needs_account":71,"seo":72,"blocks":76,"takeaways":169,"next_step":174},"is-web-scraping-legal","Is it legal?","Three questions hiding inside one","\"Is scraping legal\" bundles access, copying and use into a single question. Separating them is most of the work.","10 min",false,{"title":73,"description":74,"keywords":75},"Is Web Scraping Legal? The Three Questions That Matter","Access, copying and use are three separate legal questions that get collapsed into one. How to tell them apart and why most scraping disputes turn on only one of them.","is web scraping legal, web scraping law, scraping legality, data collection compliance, database rights",[77,82,102,108,114,119,131,155,164],{"type":78,"paragraphs":79},"prose",[80,81],"Nobody can answer \"is web scraping legal\" because it is not one question. It is at least three, they are governed by different bodies of law, and a project can be comfortably fine on two of them and in trouble on the third.","Pulling them apart is not a lawyerly evasion — it is the thing that turns an unanswerable worry into a set of specific, checkable facts.",{"type":83,"title":84,"headers":85,"rows":89},"table","The three questions",[86,87,88],"Question","What it is about","What usually decides it",[90,94,98],[91,92,93],"Access","Were you allowed to request the page at all?","Logins, access controls, and whether you circumvented anything",[95,96,97],"Copying","May you store and reproduce what came back?","Copyright in the content, and database rights in the collection",[99,100,101],"Use","May you do the thing you intend with it?","Personal data rules, competition law, contract terms",{"type":78,"title":103,"paragraphs":104},"Access is where the sharpest line sits",[105,106,107],"Requesting a page that any member of the public can load without signing in is, in most jurisdictions, close to the uncontroversial end. You are doing what a browser does, faster.","Everything changes at the login. Once there is an account, there is an agreement you accepted, and there is an access control you are operating inside rather than outside. That shifts the question from \"did you read a public page\" to \"did you comply with the terms you agreed to\", which is a contract question with a much clearer answer and much clearer consequences.","This single distinction — public page versus authenticated session — does more work than any other idea in this course. It is why lesson two spends its whole length on it.",{"type":78,"title":109,"paragraphs":110},"Copying is about the collection more than the item",[111,112,113],"A price is a fact, and facts are not generally protected by copyright. That is why price monitoring sits on firmer ground than people assume.","Two things change that. Creative content — product photography, written descriptions, reviews — is somebody's work and copying it wholesale is a different proposition from recording a number. And in the EU and UK there is a separate database right protecting substantial investment in compiling a collection, which can apply even where no individual item is protected. Taking a substantial part of someone's assembled catalogue is a different act from checking the price of forty products.","The practical consequence is that scope matters. Collecting the fields you need for a specific purpose is a narrower and more defensible act than mirroring a catalogue because you could.",{"type":115,"variant":116,"title":117,"text":118},"callout","note","Purpose changes the answer more than technique does","The same requests, producing the same rows, sit very differently depending on what happens next. Monitoring competitor prices to set your own is ordinary commercial activity that long predates the web. Reproducing a competitor's catalogue as your own storefront is not, and no amount of care in how you fetched the pages makes it so. When you write the project down, lead with the purpose.",{"type":120,"title":121,"intro":122,"items":123},"list","The factors that move a project towards the comfortable end","None is decisive alone. Together they are what a defensible position looks like.",[124,125,126,127,128,129,130],"No login, no access control, nothing circumvented","Facts rather than creative content — prices, availability, specifications","A narrow, stated purpose, and only the fields that purpose needs","A request rate that imposes no meaningful cost on the source","No personal data, or a deliberate decision about it (lesson three)","An identifiable user agent and a contact address","A written record of all of the above, made before you started",{"type":83,"title":132,"intro":133,"headers":134,"rows":137},"Worked example: one project, split into the three questions","A pricing team wants to track 1,200 of its own SKUs across five competitor sites. Written as one question it sounds unanswerable. Split into access, copying and use, every row has a plain answer and the one row that needs a decision becomes obvious.",[86,135,136],"What it actually asks here","Answer for this project",[138,141,144,147,151],[91,139,140],"Are we getting the pages the way an ordinary visitor does, without an account, without a password, without stepping around a block?","Yes. Public category and product pages, no login, standard requests, no bypass of any gate.",[95,142,143],"How much of each page do we keep, and could the stored collection substitute for the source?","Eight fields per product. No descriptions, no images, no reviews. Nobody could shop from our table.",[99,145,146],"What do we do with it, and does that compete with the source's own use of it?","Internal repricing input. Not republished, not resold, not shown to customers.",[148,149,150],"Retention","How long do we keep it, and do we still need the oldest rows?","13 months of daily snapshots, then aggregate to weekly. Nothing beyond that.",[152,153,154],"The row needing a decision","Two of the five sites require an account to see trade prices.","Those two go to the legal brief separately. The other three do not need one.",{"type":120,"title":156,"intro":157,"items":158},"What usually goes wrong","Almost every uncomfortable scraping project got there by one of these, not by a surprise in the law.",[159,160,161,162,163],"Asking \"is scraping legal\" as a single question, getting a shrug, and treating the shrug as permission.","Letting the field list grow quietly. Eight columns is a price feed. Add descriptions, images and reviews and you have built a copy of the catalogue, which is a different conversation.","Starting with the hardest site. The two login-gated competitors drag the other three into a review they never needed.","Assuming the technique decides the answer. The same HTTP request is unremarkable for internal price comparison and awkward for a public mirror of someone's catalogue.","Never writing the three answers down, so when someone asks six months later nobody can reconstruct what was decided or when.",{"type":78,"title":165,"paragraphs":166},"The honest summary",[167,168],"Collecting publicly available factual data, at a considerate rate, for your own analysis, is routine and widespread. Enormous parts of the modern web — search engines, price comparison, academic research, archiving — depend on it being so.","The cases that go wrong cluster tightly, and the clusters are recognisable: data behind a login, personal data at scale, volumes that cost the source real money, and republishing someone's collection as your own. If your project is in none of those clusters, you are in ordinary territory. If it is in one of them, you need advice specific to your jurisdiction before you start, not after.",[170,171,172,173],"Access, copying and use are three separate questions with three different bodies of law behind them.","The public-page versus logged-in line is the sharpest one and does the most work.","Facts are weakly protected; creative content and assembled databases are not. Scope narrowly.","The problem cases cluster: logins, personal data at scale, high cost to the source, republishing a catalogue.",{"text":175,"label":176,"url":177},"Next: the line that matters most — what a terms-of-service page does, and what changes at the login.","Lesson 2: public data, terms of service and logins","/learn/web-scraping-legal-and-ethical/public-data-terms-of-service-and-logins",{"slug":179,"nav_title":180,"title":181,"summary":182,"time":183,"needs_account":71,"seo":184,"blocks":188,"takeaways":284,"next_step":289},"public-data-terms-of-service-and-logins","Terms and logins","Public data, terms of service, and what changes at the login","Why a terms page you never agreed to is weaker than people think, why the one you did agree to is stronger, and where that leaves mobile app APIs.","11 min",{"title":185,"description":186,"keywords":187},"Terms of Service and Web Scraping: Browsewrap vs Clickwrap","Why terms you never clicked are weaker than terms you accepted, what changes the moment you create an account, and how that applies to mobile app APIs and consent walls.","terms of service scraping, browsewrap clickwrap, scraping behind login, public data scraping, mobile app api scraping, scraping account terms",[189,192,211,217,221,227,236,270,278],{"type":78,"paragraphs":190},[191],"Nearly every website has a terms page, nearly all of them say something about automated access, and almost nobody has read the one for the site they are about to collect from. The useful question is not what it says — it is whether you ever agreed to it.",{"type":83,"title":193,"headers":194,"rows":198},"Three levels of agreement",[195,196,197],"Situation","How strong is the agreement","Practical consequence",[199,203,207],[200,201,202],"Terms linked in a footer, never clicked","Weakest — often called browsewrap","Contested. Not nothing, but not a clear contract either.",[204,205,206],"A consent or cookie banner you dismissed","Depends entirely on wording and prominence","Grey. Worth reading what you clicked.",[208,209,210],"An account you created, accepting terms","Strongest — clickwrap, a real agreement","You are bound. If it forbids automated access, that is the answer.",{"type":78,"title":212,"paragraphs":213},"Why the distinction is not a loophole",[214,215,216],"It would be convenient to read the first row as \"footer terms don't count\". That is not the point and it is not safe.","The point is that the three rows carry genuinely different weight, and a decision that treats them as identical is making the wrong trade in both directions — either refusing a project that is perfectly ordinary, or walking into a clear contractual breach because \"nobody reads terms\".","A practical middle position: read the automated access clause before you start, record what it said and when you read it, and treat an explicit prohibition as a reason to stop and ask rather than a formality to route around. The record is what turns a judgement call into a documented decision, and documented decisions are what survive a review.",{"type":115,"variant":218,"title":219,"text":220},"warning","Creating an account is a decision, not a convenience","The moment you sign up to see prices more easily, you have accepted terms, you are inside an access control, and your activity is attributable to an identity. If those terms prohibit automated access — and most do — then automating from that account is a breach regardless of how public the underlying pages look. The pragmatic rule we operate under: if the data requires a login, we do not collect it.",{"type":78,"title":222,"paragraphs":223},"Mobile app APIs are not a shortcut around this",[224,225,226],"A recurring discovery is that a site which is hostile to scraping has a mobile app talking to a clean, fast, unprotected JSON endpoint. It is genuinely tempting — better data, less work, no HTML parsing.","It also usually sits worse than the website, not better. The app has its own terms, which you accepted on installation. The endpoint is frequently authenticated, which puts you inside an access control. And obtaining the credentials often involves inspecting traffic in a way that is itself covered by those terms.","There are cases where an app API is openly documented and unauthenticated, and those are fine. The failure mode is assuming that because something is technically reachable, it is in the same category as a public web page. It usually is not, and \"it was easier\" is not a position anybody wants to defend.",{"type":120,"title":228,"intro":229,"items":230},"Questions to answer before the first request","Five minutes each, and together they are most of the due diligence anyone will ask you for.",[231,232,233,234,235],"Is this page reachable with no account, no cookie beyond a session, and nothing circumvented?","What does the terms page say about automated access, and on what date did I read it?","Is there a published API or feed that covers the same data? Using it is better in every respect.","Would a reasonable person at that company, seeing exactly what I am doing, consider it a problem?","Can I write down the purpose in one sentence that does not sound evasive?",{"type":83,"title":237,"intro":238,"headers":239,"rows":244},"Worked example: five sites, sorted by what you actually agreed to","The same crawl across five competitors lands in three different places. Sorting the list this way takes about ten minutes and decides which sites need a conversation before the first request.",[240,241,242,243],"Site","How you reach the data","What you agreed to","Where that leaves it",[245,250,255,260,265],[246,247,248,249],"A","Public category pages, no account","Nothing. You never clicked anything.","Proceed. Conduct rules still apply.",[251,252,253,254],"B","Public pages, terms linked in the footer","Nothing you assented to.","Proceed, but read the terms so you know what you are choosing to ignore and can say so out loud.",[256,257,258,259],"C","Public pages behind a cookie banner you must dismiss","Still nothing contractual about data use in most framings, but note it.","Proceed. Record that the banner exists.",[261,262,263,264],"D","Trade prices visible only after creating an account","Everything in the terms you ticked at signup.","Stop. This is a contract question, not a scraping question.",[266,267,268,269],"E","A documented public API with a key","The API terms, plus a rate limit you can actually read.","Use the API. It is the cheapest and clearest of the five.",{"type":120,"title":156,"intro":271,"items":272},"The login line is the one people cross without noticing.",[273,274,275,276,277],"Someone on the team already has an account from a trade show or a test order, so the crawler quietly uses it and nobody records that the project changed category.","Treating \"the data is public once you are logged in\" as the same as public. The account is the thing that changed, not the data.","Reusing a personal account rather than a company one, which moves an individual's name onto a contract they did not read in this context.","Finding the mobile app's unauthenticated endpoint and treating it as a loophole. It is the same site, the same operator, and usually the same terms.","Never re-reading the terms. The sort above is accurate on the day you do it and silently rots after that, which is why the brief in the last lesson carries a date.",{"type":78,"title":279,"paragraphs":280},"The last question is the most useful one",[281,282,283],"The reasonable-person test is not a legal standard, but it is an excellent early-warning system, and it catches things the formal checks miss.","A competitor noticing that you track their public prices will shrug — they almost certainly track yours, and the practice predates e-commerce by decades. A company discovering that you have reconstructed their customer list, or that your collection is measurably slowing their site, will not.","If the honest answer to \"how would they feel about this\" is \"they would be angry, and they would have a point\", that is a signal worth acting on before anybody else gets involved.",[285,286,287,288],"Footer terms, dismissed banners and accepted account terms carry very different weight — do not flatten them.","Read the automated access clause, record what it said and when. The record is the deliverable.","An account turns a public-page question into a contract question. If data needs a login, the safe answer is not to collect it.","Mobile app APIs usually sit worse than the website, not better.",{"text":290,"label":291,"url":292},"Next: the rules that apply even when everything above is settled, because the data turned out to be about people.","Lesson 3: personal data and GDPR","/learn/web-scraping-legal-and-ethical/personal-data-and-gdpr",{"slug":294,"nav_title":295,"title":296,"summary":297,"time":183,"needs_account":71,"seo":298,"blocks":302,"takeaways":406,"next_step":411},"personal-data-and-gdpr","Personal data","Personal data, and why product scraping quietly becomes it","Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.",{"title":299,"description":300,"keywords":301},"Web Scraping and GDPR: When Product Data Becomes Personal Data","Why publicly available personal data is still regulated, the fields that turn a product dataset into a personal one — seller names, reviews, marketplace listings — and how to avoid the problem entirely.","web scraping gdpr, scraping personal data, gdpr public data, scraping reviews gdpr, marketplace seller data, data minimisation",[303,307,338,341,349,356,388,396],{"type":78,"paragraphs":304},[305,306],"The most common surprise in this whole area: publicly available personal data is still personal data. The fact that someone posted their name on a public page does not take it outside data protection law, and the intuition that \"public means fair game\" is simply wrong under the GDPR and its equivalents.","This matters for product scraping specifically, because personal data arrives in product datasets by accident far more often than by design.",{"type":83,"title":308,"headers":309,"rows":313},"Fields that look commercial and are not",[310,311,312],"Field","Why it is personal data","Safer option",[314,318,322,326,330,334],[315,316,317],"Marketplace seller name","Very often an individual or a sole trader","Keep a hashed seller id, or drop it",[319,320,321],"Review text and author","Written by an identifiable person, and sometimes revealing","Keep the rating count and average; drop the text",[323,324,325],"Q&A on a product page","Same — identifiable individuals","Drop entirely",[327,328,329],"Seller contact details","Directly identifying","Never collect",[331,332,333],"\"Sold by\" on a retailer listing","Corporate when it is a company, personal when it is not","Check which, or drop",[335,336,337],"Price, stock, SKU, title","Not personal data","Collect freely",{"type":115,"variant":218,"title":339,"text":340},"The accidental dataset","A price monitoring project aimed squarely at products ends up holding thousands of individual sellers' names and trading histories because \"sold by\" was one of the columns. Nobody decided to build a dataset about people. One field did it, and now the project is inside a regime it was never scoped for. This is the single most common way teams end up non-compliant without any bad intent.",{"type":78,"title":342,"paragraphs":343},"What the rules actually ask of you",[344,345,346,347,348],"Simplified considerably, and for the EU and UK regime specifically, there are four obligations that bite on collected data.","You need a lawful basis. For commercial research the usual candidate is legitimate interests, which requires you to balance your interest against the individual's rights and — importantly — to document that you did.","You have to tell people. There is a transparency obligation when you collect data about someone from a source other than them. There is a carve-out where notifying everyone would be disproportionate, but it is not automatic and it depends on you having thought about it.","You must minimise. Collect what the purpose needs and no more. This is also, conveniently, the rule that makes most of the problem disappear.","And you have to be able to respond. Individuals can ask what you hold and ask you to delete it, which in practice means you need to be able to find a named person in your dataset at all.",{"type":78,"title":350,"paragraphs":351},"The simplest answer is usually the right one",[352,353,354,355],"Do not collect it.","A price monitoring system does not need seller names to work. It does not need review text. Dropping those fields at the point of extraction — not filtering them out later, but never writing them down — moves the entire project out of the personal data regime and removes a category of risk, a category of obligation and a category of conversation.","Where a seller identity genuinely matters for the analysis, a stable hash preserves the ability to say \"this is the same seller as last week\" without holding anybody's name. That is sufficient for almost every commercial question people actually ask of this data.","It is also simply less to defend. A dataset that provably contains no personal data is the shortest possible answer to a security review, and security reviews are where these projects most often stall.",{"type":83,"title":357,"intro":358,"headers":359,"rows":364},"Worked example: auditing a product feed, column by column","A marketplace price feed looks entirely commercial until you list the columns and ask one question of each: could this, alone or combined with the rest, identify a living person? Here is the same feed before and after that audit.",[360,361,362,363],"Column","Identifies a person?","Needed for repricing?","Decision",[365,370,373,375,379,382,385],[366,367,368,369],"product_title","No","Yes, for matching","Keep",[371,367,372,369],"price, currency, in_stock","Yes",[374,367,368,369],"gtin, mpn",[376,377,367,378],"seller_name","Often yes on a marketplace, where a large share of sellers trade under their own name","Drop",[380,381,367,378],"seller_address","Yes, frequently a home address",[383,384,367,378],"review_author, review_text","Yes, and review text is a direct opinion attached to a named person",[386,387,367,378],"q_and_a_username","Yes, pseudonymous but linkable",{"type":120,"title":156,"intro":389,"items":390},"Nobody sets out to build a personal dataset. It assembles itself from columns that each looked harmless.",[391,392,393,394,395],"Scraping the whole page because it was easier than selecting fields, then discovering a year later that the warehouse holds seller names and review text nobody ever used.","Treating pseudonyms as anonymous. A stable username plus a purchase history plus a town is identifying in practice, whatever it looks like in isolation.","Keeping seller_name \"for debugging\" and never removing it, so a temporary convenience becomes a permanent category change.","Assuming public means unregulated. Publication by the person does not remove the obligations on whoever builds a new collection from it.","Having no deletion path, so the first time someone asks what you hold about them there is no honest answer and no mechanism to act on it.",{"type":120,"title":397,"intro":398,"items":399},"When you do need personal data, the minimum discipline","Some research genuinely requires it. If that is you, these are not optional.",[400,401,402,403,404,405],"Write the lawful basis down before collecting, with the balancing reasoning, not after","Keep the narrowest field set that answers the question","Set a retention period and actually enforce it — indefinite retention is very hard to justify","Be able to locate and delete an individual's records on request","Never collect special category data — health, politics, religion, sexuality, biometrics — without specific advice","Get a real review from someone qualified. This is the part of the course where the stakes justify the fee.",[407,408,409,410],"Public personal data is still personal data. \"It was on a public page\" is not a defence.","Product datasets acquire personal data by accident — seller names, review text, Q&A. Usually one column does it.","Minimisation is the lever: never write the field down and the whole regime stops applying.","Where seller identity matters, hash it. You keep the continuity and hold nobody's name.",{"text":412,"label":413,"url":414},"Next: conduct rather than law — rate limits, robots.txt, and being the kind of client nobody has to block.","Lesson 4: rate limits, robots.txt and good conduct","/learn/web-scraping-legal-and-ethical/rate-limits-robots-and-being-a-good-citizen",{"slug":416,"nav_title":417,"title":418,"summary":419,"time":183,"needs_account":71,"seo":420,"blocks":424,"takeaways":532,"next_step":537},"rate-limits-robots-and-being-a-good-citizen","Conduct and rate limits","Rate limits, robots.txt, and being easy to live with","The conduct half. What robots.txt is for, what rate to actually use, and why identifying yourself is the most underrated decision available.",{"title":421,"description":422,"keywords":423},"Robots.txt and Scraping Rate Limits: Good Practice","What robots.txt is and is not, how to choose a request rate, why identifying your crawler beats impersonating a browser, and the conduct that keeps you unblocked.","robots txt scraping, scraping rate limit, crawl delay, scraper user agent, web scraping ethics, polite crawling",[425,429,435,438,467,473,479,489,518,527],{"type":78,"paragraphs":426},[427,428],"The previous three lessons were about what you are permitted to do. This one is about how you do it, which is separate and in day-to-day terms more consequential — because conduct is what determines whether you stay unblocked, and whether a complaint ever reaches anybody's desk.","It is also the part most within your control. You usually cannot change whether a site has a login. You can always change how often you ask.",{"type":78,"title":430,"paragraphs":431},"What robots.txt actually is",[432,433,434],"A text file at the root of a domain, in which the site operator states which paths they would prefer automated clients not to request. It is a convention from the early web, it is honoured voluntarily, and it is not a technical control — nothing stops you ignoring it.","Whether it carries legal weight varies and is genuinely unsettled. The practical case for respecting it does not depend on that. It is the only channel a site has for saying what it wants, it costs almost nothing to honour, and disregarding it is the clearest possible evidence of bad faith if anybody ever looks at your behaviour.","Read it with some judgement, though. A blanket disallow aimed at search engine indexing is a different statement from a specific disallow on a product path. If a file forbids everything and you believe your use is legitimate, the right move is to ask the operator — not to decide unilaterally that the file did not mean you.",{"type":115,"variant":116,"title":436,"text":437},"Honour the crawl-delay even though nobody does","Where robots.txt specifies a crawl-delay, it is the site telling you exactly what rate it is comfortable with. Taking them at their word costs you a slower run and buys you a position that is very hard to criticise. It is the cheapest good-faith evidence available and it is routinely ignored.",{"type":83,"title":439,"intro":440,"headers":441,"rows":446},"Choosing a rate","Rules of thumb, not standards. When in doubt, be slower — nobody has ever been blocked for being too polite.",[442,443,444,445],"Site type","Reasonable concurrency","Reasonable gap","Why",[447,452,457,462],[448,449,450,451],"Large retailer or marketplace","2 to 4","1 to 2 seconds","They serve far more than this per second already",[453,454,455,456],"Mid-sized e-commerce","1 to 2","2 to 5 seconds","Your traffic is visible in their analytics",[458,459,460,461],"Small or independent site","1","5 to 10 seconds","You could be a noticeable share of their load",[463,464,465,466],"Anything that has slowed down","Back off immediately","Exponential","A 429 or a rising latency is a request to stop",{"type":78,"title":468,"paragraphs":469},"Frequency is a separate question from rate, and it is the bigger one",[470,471,472],"Rate is how fast you go through a run. Frequency is how often you do the run, and it is where most unnecessary load comes from.","The honest test: how often does the underlying value actually change? A price that moves twice a week does not need hourly collection, and most schedules are set by what felt responsive at the time rather than by any measurement. Halving your frequency halves your load on the source, halves your cost, and in most cases changes nothing anybody downstream would notice.","Better still, make the frequency match the volatility. Fast-moving categories daily, stable ones weekly. This is both more considerate and cheaper, which is a rare combination.",{"type":78,"title":474,"paragraphs":475},"Identify yourself",[476,477,478],"The default instinct is to blend in — copy a browser's user agent and hope nobody looks. It is understandable and it is usually the wrong call.","A user agent naming your organisation with a contact URL means that when somebody does look at their logs, they find a named, contactable, well-behaved client. The realistic outcomes are an allowlist, an email asking you to slow down, or nothing at all. The outcomes available to an anonymous client failing a fingerprint check are a ban and, if it ever escalates, a much worse story.","There is a real tension here, because some protection layers are more hostile to a declared crawler than to a convincing browser impersonation. Our position is that the long game favours being identifiable: impersonation is an arms race you lose eventually, and a relationship is durable in a way that a working header set is not.",{"type":120,"title":480,"items":481},"Conduct that keeps you welcome",[482,483,484,485,486,487,488],"Back off on 429 and 503 rather than retrying immediately — those are explicit requests to slow down","Run heavy jobs outside the source's peak hours","Cache, and never request a page twice when once would do","Request only what you need — no images, no assets, no pages outside the scope","Publish a contact address and answer it","Stop promptly if asked, and talk to them before resuming","Review the schedule periodically. Jobs set up two years ago are usually running far more often than anyone currently needs.",{"type":83,"title":490,"intro":491,"headers":492,"rows":497},"Worked example: the same 38,600 pages at four different rates","One competitor, 38,600 product URLs, run once a day. The rate you pick decides both how long the run takes and how visible you are in their logs. Their own traffic is the number that matters: a mid-size retailer serving roughly 40,000 page views a day is handling about 0.5 requests per second on average.",[493,494,495,496],"Rate","Run duration","Share of their average traffic","How it reads in their logs",[498,503,508,513],[499,500,501,502],"10 req/s","64 minutes","About 20x their own average rate","A spike. Someone will look, and they should.",[504,505,506,507],"2 req/s","5 hours 22 minutes","About 4x","Still the loudest client they have that hour.",[509,510,511,512],"0.5 req/s","21 hours 26 minutes","About 1x","Indistinguishable from a busy customer, but it no longer fits in a day.",[514,515,516,517],"1 req/s, split across 2 nightly windows","10 hours 43 minutes, off-peak","About 2x, during their quietest hours","Invisible in practice. This is the one to pick.",{"type":120,"title":156,"intro":519,"items":520},"Most of the damage comes from the schedule rather than the rate.",[521,522,523,524,525,526],"Tuning the rate carefully and then running six scrapers in parallel against the same host, so the real figure is six times the one on the dashboard.","Running hourly because the scheduler made it easy, when prices on that site change twice a week. Frequency multiplies everything.","Retrying failures immediately, so the moment a site is struggling is exactly the moment your traffic triples.","Running at the top of the hour like everyone else's cron, which stacks your load onto theirs.","Using a generic browser user agent with no contact address, which turns a solvable conversation into a block. Being identifiable is what lets someone email you instead of banning you.","Scraping the same page five times to get five fields, because the extraction was written field by field rather than page by page.",{"type":78,"title":528,"paragraphs":529},"The underlying principle",[530,531],"Be the kind of client a site operator would not bother blocking, because blocking you would be more effort than tolerating you.","That is not only an ethical stance, it is the most reliable operational strategy available. The scrapers that get blocked are the ones that are expensive, anonymous and relentless. The ones that run for years are polite, identifiable and ask for exactly what they need.",[533,534,535,536],"robots.txt is a convention, not a control, and honouring it costs almost nothing while disregarding it is clear evidence of bad faith.","Be slower than you think you need. Nobody has ever been blocked for politeness.","Frequency matters more than rate — match collection frequency to how often the value actually changes.","Identify yourself with a contact address. Impersonation is an arms race you lose; a relationship is durable.",{"text":538,"label":539,"url":540},"Finally: turning all of this into the one page that gets you an answer instead of a reflexive no.","Lesson 5: what to put in front of your legal team","/learn/web-scraping-legal-and-ethical/what-to-put-in-front-of-your-legal-team",{"slug":542,"nav_title":543,"title":544,"summary":545,"time":546,"needs_account":71,"seo":547,"blocks":551,"takeaways":672,"next_step":677},"what-to-put-in-front-of-your-legal-team","Briefing legal","What to put in front of your legal team","A one-page brief that gets a real answer, the three mistakes that guarantee a no, and the position we operate under ourselves.","12 min",{"title":548,"description":549,"keywords":550},"How to Get Legal Approval for a Web Scraping Project","A one-page brief template for data collection projects, the three framings that guarantee a reflexive no, and an example acceptable-use position.","web scraping legal approval, data collection compliance brief, legitimate interests assessment, scraping policy, acceptable use data collection",[552,556,585,588,607,613,619,651,660],{"type":78,"paragraphs":553},[554,555],"Most data collection projects that die in legal review do not die because the answer was no. They die because the question was unanswerable, so the safe response was no.","\"Are we allowed to scrape competitor websites?\" has no good answer. A lawyer hearing it has to imagine the worst version of what you might mean and advise against that. One page of specifics changes the entire conversation, and writing it takes about an hour.",{"type":557,"title":558,"intro":559,"items":560},"steps","The one-page brief","Eight sections, a few sentences each. Written before you start, not after someone asks.",[561,564,567,570,573,576,579,582],{"title":562,"text":563},"Purpose, in one sentence","What business decision this data informs. \"Weekly price position against eleven named competitors on the four hundred SKUs we actively compete on.\" Specific, bounded, obviously ordinary.",{"title":565,"text":566},"Sources, named","The actual domains. Not \"competitor websites\" — the list. Vagueness here reads as evasion even when it is laziness.",{"title":568,"text":569},"Fields, exhaustively","Every column you will store. This is where you demonstrate minimisation, and where a reviewer can see at a glance that there is no personal data in it.",{"title":571,"text":572},"Access method","Public pages, no login, nothing circumvented. State it plainly, because it is the question they most want answered.",{"title":574,"text":575},"Terms review","What each source's terms say about automated access, and the date you read them. A table with one row per source.",{"title":577,"text":578},"Rate and frequency","Requests per minute, concurrency, how often the run happens, and how that compares to the source's own traffic. Numbers, not adjectives.",{"title":580,"text":581},"Personal data position","Either \"none collected, here is the field list\" — much the better answer — or your lawful basis and retention period.",{"title":583,"text":584},"Who sees the output","Internal only, a specific team, or something customer-facing. Republishing changes the analysis completely and is the thing they will ask about if you do not say.",{"type":115,"variant":116,"title":586,"text":587},"The date on the terms review is doing real work","It shows a process rather than an assumption, and it gives you a defensible position if the terms change later: you checked, on a date, and recorded what they said. A reviewer reading that knows they are dealing with someone who will also notice when something changes.",{"type":83,"title":589,"headers":590,"rows":594},"Three framings that guarantee a no",[591,592,593],"What people say","What it sounds like","Say this instead",[595,599,603],[596,597,598],"\"It's all public data\"","You have not thought about personal data or terms","\"No login, no personal data, here is the field list\"",[600,601,602],"\"Everyone does it\"","You have no position of your own","\"This is standard market research; here is our conduct policy\"",[604,605,606],"\"They can't tell it's us\"","You are relying on not being caught","\"We identify ourselves and publish a contact address\"",{"type":78,"title":608,"paragraphs":609},"Expect conditions, not a verdict",[610,611,612],"A good review rarely produces a clean yes. It produces a yes with conditions, and the conditions are usually reasonable and cheap: drop two fields, halve the frequency, exclude one source whose terms are explicit, add a retention limit, re-check the terms annually.","Treat that as the success case. A conditional approval is a documented, bounded authorisation that protects you and the project, and it is far more valuable than an unconditional shrug that nobody wrote down.","If the answer is a flat no, ask which of the eight sections caused it. Frequently it is one — a single source, or one field — and the project survives without it.",{"type":78,"title":614,"paragraphs":615},"The position we operate under",[616,617,618],"Worth stating, because a vendor who will not tell you where their line is has not thought about it, and you should ask every vendor this question.","We collect from public pages only. No logins, no credential sharing, no circumvention of access controls. We rate-limit by default and back off on 429 and 503. We do not build datasets about individuals — no seller names, no review text, no contact details. We honour removal requests from site operators. And we tell customers plainly when a source cannot be collected rather than filling the gap with a plausible number, because a coverage gap is a known quantity and an invented row is not.","That last one is a data integrity commitment rather than a legal one, but it belongs in the same list. The projects that cause problems downstream are rarely the ones with a documented gap. They are the ones where somebody decided a gap looked bad.",{"type":83,"title":620,"intro":621,"headers":622,"rows":626},"Worked example: the same project, briefed badly and briefed well","Two versions of one request about the same five competitors. The left column is what legal teams usually receive. The right column is the same project described in terms somebody can actually sign off, line for line.",[623,624,625],"What they usually get","What they can answer","Why the difference matters",[627,631,635,639,643,647],[628,629,630],"\"Can we scrape competitor websites?\"","\"We want to read 1,200 public product pages across five named sites, once a day, off-peak, at roughly 1 request per second.\"","The first has no boundary, so the safe answer is no. The second has five facts to check.",[632,633,634],"\"We'd collect pricing data.\"","\"Eight fields: title, GTIN, MPN, price, currency, availability, shipping cost, URL. No descriptions, no images, no reviews, no seller names.\"","A named field list is the single strongest thing in the brief. It proves the collection cannot substitute for the source.",[636,637,638],"\"For analysis.\"","\"Input to internal repricing. Not republished, not resold, not shown to customers or in marketing.\"","Use decides more of the answer than technique does.",[640,641,642],"\"It's all public.\"","\"Three sites need no account. Two show trade prices only behind a login, and we have excluded those two pending your view.\"","Flagging the hard case yourself is what makes the rest credible.",[644,645,646],"(no date)","\"Terms of each site reviewed on 14 March. We re-review every six months and on any redesign.\"","An undated review is an assertion. A dated one is a control.",[648,649,650],"(no exit)","\"We stop within one business day of any request from the site operator, and here is the mailbox that receives it.\"","Reversibility turns a permanent decision into a revocable one.",{"type":120,"title":156,"intro":652,"items":653},"A no from legal is usually a response to the brief, not to the project.",[654,655,656,657,658,659],"Asking the abstract question. \"Is scraping legal\" has no answer that helps anyone, and the only safe response to an unbounded question is no.","Asking after the crawler is already running, which turns an approval into an incident review.","Hiding the login-gated sites in the middle of the list instead of naming them as the open question.","Promising a field list and then letting it grow, so the thing that was approved and the thing that runs drift apart within a quarter.","Treating the answer as permanent. Conditions expire, sites get redesigned, terms change, and a review with no date on it stops being evidence of anything.","Leaving no owner. If no named person re-reviews on a schedule, the brief is a document rather than a control.",{"type":120,"title":661,"intro":662,"items":663},"The closing checklist","If you can tick all of these, you are in ordinary territory and you have the paperwork to show it.",[664,665,666,667,668,669,670,671],"Purpose written in one sentence, and it does not sound evasive","Public pages only, no account, nothing circumvented","Field list is the minimum the purpose needs, and contains no personal data","Terms reviewed per source, with dates recorded","Rate and frequency stated as numbers, and justified against how often the data changes","Identifiable user agent with a working contact address","Someone qualified has read the brief and signed off, with any conditions written down","A review date in the calendar — terms change, catalogues change, and so does your purpose",[673,674,675,676],"Projects fail review because the question was unanswerable, not because the answer was no.","Eight sections, one page, written before you start. Named sources, exhaustive field list, numbers for rate and frequency.","\"It's public\", \"everyone does it\" and \"they can't tell it's us\" each guarantee a no. Replace each with a specific.","Expect conditional approval and treat it as the win. Ask every vendor where their own line is.",{"text":678,"label":679,"url":24},"That is the end of the course. If you want to see what collected data actually looks like before committing to anything, every retailer page publishes real measured output — including the ones where we say plainly that nothing readable came back.","See real run output by retailer",[681,734,773,812,851,889],{"order":682,"slug":683,"title":684,"subtitle":685,"cardText":686,"level":687,"time":688,"lessonCount":689,"lessons":690},1,"competitor-price-monitoring","Build a competitor price monitoring pipeline","Price monitoring looks like a scraping problem for about a week. Then you discover that scraping was the easy part, and the project actually lives or dies on which competitors you picked, whether their listings are really the same product as yours, and whether anyone notices the morning the feed comes back half empty. This course is those eight decisions, in the order you have to make them.","From \"we check three competitors by hand on Mondays\" to a feed you trust enough to reprice from. The eight decisions in order, including the two that quietly ruin most projects.","No coding required","8 lessons, about 90 minutes",8,[691,697,702,708,713,719,724,729],{"slug":692,"navTitle":693,"title":694,"summary":695,"time":696,"needsAccount":71},"what-is-competitor-price-monitoring","What it actually is","What competitor price monitoring actually is","The four stages of a price pipeline, why only two of them are scraping, and the one question to ask before you build anything.","9 min",{"slug":698,"navTitle":699,"title":700,"summary":701,"time":183,"needsAccount":71},"choose-competitors-and-skus","Choosing what to track","Choosing which competitors and which SKUs to track","How to build a list that is small enough to afford and large enough to matter, using margin at risk rather than gut feel.",{"slug":703,"navTitle":704,"title":705,"summary":706,"time":546,"needsAccount":707},"find-competitor-product-urls","Finding product URLs","Finding every competitor product URL without copying them by hand","Four ways to get a competitor's full product URL list, ranked by how much work they are, and what to do when none of them work.",true,{"slug":709,"navTitle":710,"title":711,"summary":712,"time":546,"needsAccount":707},"extract-price-stock-and-shipping","Extracting the fields","Getting price, stock and shipping off the page","Which fields to extract, why the sale price is two fields and not one, and the four ways a price appears on a page.",{"slug":714,"navTitle":715,"title":716,"summary":717,"time":718,"needsAccount":71},"match-listings-to-your-catalogue","Matching to your catalogue","Matching competitor listings to your own catalogue","The stage that decides whether your feed is intelligence or fiction, and the denominator trick that makes bad match rates look good.","13 min",{"slug":720,"navTitle":721,"title":722,"summary":723,"time":183,"needsAccount":71},"schedule-runs-and-catch-silent-failure","Scheduling and data quality","Scheduling runs and catching silent data loss","How often to actually check, and the four alerts that catch a degrading feed before someone reprices from it.",{"slug":725,"navTitle":726,"title":727,"summary":728,"time":696,"needsAccount":707},"export-to-sheets-bi-and-erp","Getting the data out","Getting the data into Sheets, BI or your ERP","Four delivery routes ranked by how likely they are to actually get used, and the column contract that stops downstream jobs breaking.",{"slug":730,"navTitle":731,"title":732,"summary":733,"time":546,"needsAccount":71},"turn-price-data-into-repricing-rules","From data to decisions","Turning price data into repricing decisions","Why \"match the cheapest\" destroys margin, what a rule needs besides a competitor price, and how to start without automating anything.",{"order":735,"slug":736,"title":737,"subtitle":738,"cardText":739,"level":740,"time":741,"lessonCount":5,"lessons":742},2,"ai-agent-web-data-mcp","Give your AI agent live web data via MCP","Ask an assistant what a product costs today and you will usually get a number. It is often wrong, and it is always wrong in the same way: the model is reconstructing a plausible price from training data rather than looking at a page. This course is about closing that gap properly — what the Model Context Protocol actually is, how to wire a server into a client, how to design tools a model can use without hand-holding, and what to put in place before an agent spends your money.","Your agent is confidently wrong about prices because it has never seen one. What MCP is, how to connect a server, how to design tools a model can actually use, and the guardrails you need before you let it loose.","Comfortable editing a config file","6 lessons, about 60 minutes",[743,748,753,758,763,768],{"slug":744,"navTitle":745,"title":746,"summary":747,"time":696,"needsAccount":71},"what-is-mcp","What MCP is","What MCP actually is, in plain terms","The Model Context Protocol described without jargon: what problem it solves, its three primitives, and when it is the wrong tool.",{"slug":749,"navTitle":750,"title":751,"summary":752,"time":70,"needsAccount":71},"why-agents-get-live-data-wrong","Why agents get it wrong","Why your agent's answer about a price is wrong","Four distinct failure modes that all look identical from the outside, and how to tell which one you have before you try to fix it.",{"slug":754,"navTitle":755,"title":756,"summary":757,"time":70,"needsAccount":71},"connect-an-mcp-server","Connecting a server","Connecting an MCP server and proving it works","The config for local and remote servers, the four things that go wrong, and how to verify the tools registered rather than assuming.",{"slug":759,"navTitle":760,"title":761,"summary":762,"time":183,"needsAccount":71},"design-tools-an-agent-can-use","Designing usable tools","Designing tools an agent can actually use","A connected server is not a useful server. The model only sees your tool names, descriptions and parameter schemas, so those three things are the entire user interface. Here is what makes a tool get called correctly and what makes it get ignored.",{"slug":764,"navTitle":765,"title":766,"summary":767,"time":546,"needsAccount":707},"give-an-agent-a-scraper","Giving an agent a scraper","Giving an agent a real price feed","A worked example. Connect the ScrapeWise MCP server to a client, let the agent read a live scraper's output, and watch where the hand-off between \"the data is right\" and \"the answer is right\" actually breaks.",{"slug":769,"navTitle":770,"title":771,"summary":772,"time":183,"needsAccount":71},"guardrails-cost-and-untrusted-content","Guardrails and cost","Guardrails, cost control and untrusted content","Live web access turns an agent into something that can spend money and read text written by strangers. Neither is a reason not to do it. Both are reasons to put limits in before you need them.",{"order":774,"slug":775,"title":776,"subtitle":777,"cardText":778,"level":779,"time":780,"lessonCount":5,"lessons":781},3,"product-data-api","Pull product data over an API","Search volume for \"\u003Cretailer> API documentation\" is enormous and the documentation mostly does not exist. Amazon, Walmart, Target, Home Depot — developers keep looking for a product endpoint that was never published, or that was published and then locked behind a partner agreement. So you end up calling a web data API instead: something that takes a URL and gives you back the fields. This course is about doing that properly, from the first authenticated request to a feed your warehouse can depend on.","Every retailer gets asked for an API and most of them never ship one, so you end up calling somebody else's. What a product data API actually returns, how to declare the fields you want, why long runs are asynchronous, and how to retry without paying twice.","Comfortable with HTTP and JSON","6 lessons, about 70 minutes",[782,787,792,797,802,807],{"slug":783,"navTitle":784,"title":785,"summary":786,"time":70,"needsAccount":71},"when-an-api-beats-a-scraper","API, scraper or dataset","When an API beats writing your own scraper","Three ways to get product data, the honest cost of each, and the specific question that decides between them.",{"slug":788,"navTitle":789,"title":790,"summary":791,"time":70,"needsAccount":71},"authentication-and-your-first-call","Auth and the first call","Authentication, keys, and your first real request","Bearer tokens versus query-string keys, where to keep the secret, and how to read the first response you get back.",{"slug":793,"navTitle":794,"title":795,"summary":796,"time":546,"needsAccount":707},"declare-the-fields-you-want","Declaring the fields","Declaring a schema, and why your fields came back empty","An extractor returns what you asked for, and most people ask badly. How to declare fields, why types matter, and the one mistake that silently drops a column.",{"slug":798,"navTitle":799,"title":800,"summary":801,"time":183,"needsAccount":71},"asynchronous-runs-and-polling","Async runs and polling","Asynchronous runs, polling, and partial results","Why collection APIs hand back a job rather than data, how to poll without hammering, and what to do with a run that finished eighty per cent done.",{"slug":803,"navTitle":804,"title":805,"summary":806,"time":183,"needsAccount":71},"errors-retries-and-double-billing","Errors and retries","Errors, retries, and not paying twice","Which failures are worth retrying, how idempotency keys stop a retry becoming a second invoice, and the error class that means stop rather than try harder.",{"slug":808,"navTitle":809,"title":810,"summary":811,"time":546,"needsAccount":707},"put-the-feed-into-your-stack","Into your stack","Putting the feed into your stack without it drifting","Scheduling, loading, and the schema decisions that determine whether a price feed is still trustworthy in six months.",{"order":813,"slug":814,"title":815,"subtitle":816,"cardText":817,"level":818,"time":819,"lessonCount":5,"lessons":820},4,"keep-scrapers-alive","Keep scrapers alive after the first week","Writing a scraper is a pleasant afternoon. Keeping forty of them returning correct data for two years is a different discipline, and almost nothing written about scraping covers it. This course is the maintenance half: how pages fail, how to tell a block from a redesign from an empty result, what makes a selector durable, and how to find out your feed is wrong before the person using it does.","Every scraper works on the day you write it. This course is about the other three hundred and sixty four days: why they break, how to read a failure instead of guessing at it, which selectors survive a redesign, and how to notice a feed has gone quietly wrong before somebody prices against it.","You already have something running","6 lessons, about 65 minutes",[821,826,831,836,841,846],{"slug":822,"navTitle":823,"title":824,"summary":825,"time":70,"needsAccount":71},"why-scrapers-break","Why scrapers break","The five reasons a scraper stops working","Breakage is not one problem. It is five, they have different fixes, and treating them as one is why maintenance feels endless.",{"slug":827,"navTitle":828,"title":829,"summary":830,"time":546,"needsAccount":71},"read-the-failure-not-the-symptom","Read the failure","Read the failure, not the symptom","A diagnosis routine that gets you to the cause in ten minutes, and the three false conclusions it is designed to prevent.",{"slug":832,"navTitle":833,"title":834,"summary":835,"time":183,"needsAccount":71},"selectors-that-survive-a-redesign","Durable selectors","Selectors that survive a redesign","A ranking of extraction targets by how long they last, why generated class names are a trap, and the fallback chain worth building.",{"slug":837,"navTitle":838,"title":839,"summary":840,"time":183,"needsAccount":71},"bot-walls-and-what-actually-works","Bot walls","Bot walls, and what actually changes the outcome","What a protection layer is measuring, why the laptop test lies to you, and the boring answers that work better than the clever ones.",{"slug":842,"navTitle":843,"title":844,"summary":845,"time":183,"needsAccount":71},"monitor-the-feed-not-the-run","Monitor the feed","Monitor the feed, not the run","Six checks that catch a scraper that is lying to you, and how to set thresholds that do not train everyone to ignore the alert.",{"slug":847,"navTitle":848,"title":849,"summary":850,"time":70,"needsAccount":71},"decide-what-to-do-when-a-site-wins","When a site wins","Deciding what to do when a site wins","A decision rule for fix, work around, or stop — and how to report a coverage gap so that it is useful rather than an apology.",{"order":852,"slug":853,"title":854,"subtitle":855,"cardText":856,"level":857,"time":780,"lessonCount":5,"lessons":858},5,"matching-products-across-sites","Match the same product across different sites","A price comparison is a claim that two things are the same thing. Almost every disappointing price monitoring project fails here rather than at collection: the prices were fine and the matches were not. This course is about doing the matching properly — leaning on identifiers where they exist, being honest about confidence where they do not, and measuring the result in a way that does not flatter you.","Collecting prices is the easy half. Deciding that this product on your site and that product on a competitor's are the same thing is where price monitoring actually succeeds or fails. Identifiers, fuzzy matching, variants, confidence scores and how to measure your match rate without flattering yourself.","You have data from more than one site",[859,864,869,874,879,884],{"slug":860,"navTitle":861,"title":862,"summary":863,"time":70,"needsAccount":71},"why-matching-is-the-hard-part","Why matching is hard","Why matching is the hard part","The same object is described differently by every retailer that sells it, and the differences are not noise — they are deliberate.",{"slug":865,"navTitle":866,"title":867,"summary":868,"time":183,"needsAccount":71},"identifiers-first-gtin-ean-mpn","Identifiers first","Identifiers first: GTIN, EAN, UPC and MPN","What each identifier means, how to validate one before trusting it, and the three ways a correct-looking barcode still produces a wrong match.",{"slug":870,"navTitle":871,"title":872,"summary":873,"time":546,"needsAccount":71},"when-there-is-no-barcode","No barcode","Matching when there is no barcode","Normalisation, blocking, scoring on multiple signals, and why the string similarity algorithm matters far less than everyone assumes.",{"slug":875,"navTitle":876,"title":877,"summary":878,"time":546,"needsAccount":71},"variants-bundles-and-multipacks","Variants and packs","Variants, bundles and multipacks","The highest-scoring wrong matches all live here. Normalising to a comparable unit, and knowing when two things are genuinely not comparable.",{"slug":880,"navTitle":881,"title":882,"summary":883,"time":183,"needsAccount":707},"score-confidence-and-build-a-review-queue","Confidence and review","Confidence scores and a review queue worth using","Why one score is not enough, how to set the two thresholds, and how to order a queue so an hour of human attention is worth having.",{"slug":885,"navTitle":886,"title":887,"summary":888,"time":183,"needsAccount":71},"measure-your-match-rate-honestly","Measure it honestly","Measure your match rate honestly","The denominator everyone picks is the flattering one. Precision, recall, a hand-labelled sample, and what to do with a number you do not like.",{"order":5,"slug":4,"title":17,"subtitle":18,"cardText":8,"level":6,"time":7,"lessonCount":852,"lessons":890},[891,892,893,894,895],{"slug":66,"navTitle":67,"title":68,"summary":69,"time":70,"needsAccount":71},{"slug":179,"navTitle":180,"title":181,"summary":182,"time":183,"needsAccount":71},{"slug":294,"navTitle":295,"title":296,"summary":297,"time":183,"needsAccount":71},{"slug":416,"navTitle":417,"title":418,"summary":419,"time":183,"needsAccount":71},{"slug":542,"navTitle":543,"title":544,"summary":545,"time":546,"needsAccount":71},1791047866714]