[{"data":1,"prerenderedAt":70},["ShallowReactive",2],{"$f8L8qZ-g7tZqmb6Gc4tKa50HxAiWNOK98rzqDWzWpPho":3},{"title":4,"date":5,"dateModified":6,"datePublished":7,"dateModifiedISO":7,"image":8,"content":9,"faq":6,"metaTitle":10,"metaDescription":11,"author":12,"authorBio":6,"authorLinkedin":6,"authorTitle":6,"authorPhoto":6,"lastReviewed":6,"researchBasis":6,"category":13,"readingTime":14,"related":15,"prev":33,"next":6,"toc":36,"takeaways":69},"HTTP 200 Doesn't Mean You Got the Page: 8 European Marketplaces Measured","30 September 2026",null,"2026-09-30","/img/news/http-200-not-success-eu-marketplaces-2026.png","\u003Cp>Your scraper logs a 100% success rate. Your dashboard is green. Your price table is empty, or worse, quietly wrong.\u003C/p>\n\u003Cp>This is the most expensive failure mode in production scraping, and it is not caused by a block. It is caused by trusting the wrong signal. A status code tells you the server answered. It tells you nothing about whether it answered with the page you asked for.\u003C/p>\n\u003Cp>In late September 2026 we fetched eight European marketplaces and retailers with a single plain HTTP request each — no browser, no proxy, just a realistic User-Agent — and read every byte that came back. Four of the eight answered \u003Ccode>200 OK\u003C/code>. Only two of those four actually handed over the page.\u003C/p>\n\u003Cp>\u003Cstrong>Quick answer:\u003C/strong> never gate a scrape on the HTTP status code. Half the \u003Ccode>200\u003C/code> responses we measured were empty shells: one served the homepage in response to a search URL, another served a bot-challenge payload that named its own verdict in the body. Gate on \u003Cem>content assertions\u003C/em> instead — assert a currency token, a structured-data node, or a minimum count of product blocks, and treat a failed assertion as a failed request regardless of status. The four sites that returned \u003Ccode>403\u003C/code> were the honest ones. Full measurements in the \u003Ca href=\"#what-we-measured-8-european-marketplaces-one-plain-fetch-each\">table below\u003C/a>.\u003C/p>\n\u003Ch2 id=\"why-an-http-status-code-cannot-tell-you-if-a-scrape-worked\">Why an HTTP Status Code Cannot Tell You If a Scrape Worked\u003C/h2>\n\u003Cp>A status code describes the fate of the HTTP transaction, not the usefulness of the payload. Those are different questions, and modern bot management has every incentive to decouple them.\u003C/p>\n\u003Cp>Returning \u003Ccode>403\u003C/code> to a scraper is informative. It confirms detection, tells the operator exactly which request shape got caught, and invites iteration. Returning \u003Ccode>200\u003C/code> with a stripped body is strictly better defence: the crawler records a success, moves on, and nobody investigates for weeks. The cost of the block lands on the scraper&#39;s data quality rather than its error log.\u003C/p>\n\u003Cp>So the sites that are hardest to scrape are frequently the ones that never show you an error at all. If you only ever assert \u003Ccode>response.status == 200\u003C/code>, you have built a monitor that is blind to the most sophisticated half of your targets.\u003C/p>\n\u003Caside class=\"article__usecase-card\">\u003Cdiv class=\"article__usecase-label\">Related use case\u003C/div>\u003Ch3 class=\"article__usecase-title\">Any-site data scraper\u003C/h3>\u003Cp class=\"article__usecase-blurb\">No-code extraction from any website. Managed infrastructure, no anti-bot headaches.\u003C/p>\u003Ca class=\"article__usecase-link\" href=\"/use-cases/data-scraper\">See how it works →\u003C/a>\u003C/aside>\u003Ch2 id=\"what-we-measured-8-european-marketplaces-one-plain-fetch-eac\">What We Measured: 8 European Marketplaces, One Plain Fetch Each\u003C/h2>\n\u003Cp>One plain HTTP GET per site, realistic browser User-Agent, no proxy and no JavaScript execution. Where the result was ambiguous we loaded the same URL again in a real browser to see what a human would be shown. eMAG was measured 29 September 2026; the other seven on 30 September 2026.\u003C/p>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Site\u003C/th>\n\u003Cth align=\"right\">Status\u003C/th>\n\u003Cth align=\"right\">Body\u003C/th>\n\u003Cth>What was actually in it\u003C/th>\n\u003Cth align=\"center\">Usable?\u003C/th>\n\u003C/tr>\n\u003C/thead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>\u003Ca href=\"https://scrapewise.ai/custom-scrapers/coolblue\">Coolblue\u003C/a>\u003C/td>\n\u003Ctd align=\"right\">200\u003C/td>\n\u003Ctd align=\"right\">1.4 MB\u003C/td>\n\u003Ctd>Structured price, currency, stock, brand, rating; search returned 24 products\u003C/td>\n\u003Ctd align=\"center\">\u003Cstrong>Yes\u003C/strong>\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ca href=\"https://scrapewise.ai/custom-scrapers/emag\">eMAG\u003C/a>\u003C/td>\n\u003Ctd align=\"right\">200\u003C/td>\n\u003Ctd align=\"right\">—\u003C/td>\n\u003Ctd>Product tiles, prices and stock wording already in the HTML\u003C/td>\n\u003Ctd align=\"center\">\u003Cstrong>Yes\u003C/strong>\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ca href=\"https://scrapewise.ai/custom-scrapers/ceneo\">Ceneo\u003C/a>\u003C/td>\n\u003Ctd align=\"right\">\u003Cstrong>200\u003C/strong>\u003C/td>\n\u003Ctd align=\"right\">21 KB\u003C/td>\n\u003Ctd>The Ceneo \u003Cstrong>homepage title\u003C/strong> on a search URL. Stylesheet variables. No structured data, no price token anywhere\u003C/td>\n\u003Ctd align=\"center\">No\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ca href=\"https://scrapewise.ai/custom-scrapers/cdiscount\">Cdiscount\u003C/a>\u003C/td>\n\u003Ctd align=\"right\">\u003Cstrong>200\u003C/strong>\u003C/td>\n\u003Ctd align=\"right\">13.7 KB\u003C/td>\n\u003Ctd>A bot-challenge payload naming its own verdict: request marked for a JavaScript challenge, bot category \u003Ccode>unknown\u003C/code>. No offers, no prices\u003C/td>\n\u003Ctd align=\"center\">No\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ca href=\"https://scrapewise.ai/custom-scrapers/allegro\">Allegro\u003C/a>\u003C/td>\n\u003Ctd align=\"right\">403\u003C/td>\n\u003Ctd align=\"right\">808 B\u003C/td>\n\u003Ctd>One instruction: enable JavaScript and disable your ad blocker\u003C/td>\n\u003Ctd align=\"center\">No\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ca href=\"https://scrapewise.ai/custom-scrapers/alza\">Alza\u003C/a>\u003C/td>\n\u003Ctd align=\"right\">403\u003C/td>\n\u003Ctd align=\"right\">65.7 KB\u003C/td>\n\u003Ctd>A challenge body. In a browser: Cloudflare&#39;s \u003Ccode>Just a moment...\u003C/code> interstitial\u003C/td>\n\u003Ctd align=\"center\">No\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ca href=\"https://scrapewise.ai/custom-scrapers/zalando\">Zalando\u003C/a>\u003C/td>\n\u003Ctd align=\"right\">403\u003C/td>\n\u003Ctd align=\"right\">—\u003C/td>\n\u003Ctd>Refused, including the bare homepage\u003C/td>\n\u003Ctd align=\"center\">No\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ca href=\"https://scrapewise.ai/custom-scrapers/fnac\">Fnac\u003C/a>\u003C/td>\n\u003Ctd align=\"right\">403\u003C/td>\n\u003Ctd align=\"right\">—\u003C/td>\n\u003Ctd>In a browser: a page titled \u003Ccode>FNAC DARTY - Maintenance\u003C/code>\u003C/td>\n\u003Ctd align=\"center\">No\u003C/td>\n\u003C/tr>\n\u003C/tbody>\u003C/table>\n\u003Cp>Read the two bolded rows again. Those are the dangerous ones. A scraper checking status codes recorded six failures and two successes. The reality was six failures and two successes — but it identified the wrong six, and it would have written Ceneo&#39;s and Cdiscount&#39;s empty responses into the database as good data.\u003C/p>\n\u003Ch2 id=\"the-three-shapes-of-a-block\">The Three Shapes of a Block\u003C/h2>\n\u003Ch3 id=\"1-the-honest-refusal\">1. The honest refusal\u003C/h3>\n\u003Cp>Allegro, Alza, Zalando and Fnac all returned \u003Ccode>403\u003C/code>. This is the easy case: your error handling already sees it, your retry logic already fires, your alerting already knows.\u003C/p>\n\u003Cp>Two details worth keeping. Allegro&#39;s 808-byte body is unusually explicit — it states its own requirement, which is that JavaScript must run and no ad blocker may be present. And Zalando refused the \u003Cstrong>bare homepage\u003C/strong>, not just a deep listing URL, which rules out a malformed path as the explanation. That distinction matters when you are debugging: always probe the site root before concluding your URL construction is wrong.\u003C/p>\n\u003Cp>Fnac is the one we will not overclaim. A \u003Ccode>403\u003C/code> plus a page titled \u003Ccode>FNAC DARTY - Maintenance\u003C/code> is consistent with a soft block \u003Cem>and\u003C/em> consistent with genuine maintenance. From outside the network you cannot tell the difference, and any post that tells you otherwise is guessing.\u003C/p>\n\u003Ch3 id=\"2-the-silent-substitution\">2. The silent substitution\u003C/h3>\n\u003Cp>Ceneo answered our search URL with \u003Ccode>200\u003C/code> and 21 KB — and served the \u003Cstrong>homepage\u003C/strong>. The title tag was Ceneo&#39;s root title, the body was largely stylesheet variables, and there was no structured product data and not a single price token in the entire response.\u003C/p>\n\u003Cp>Nothing about that response is an error. The connection succeeded, the server was healthy, the HTML was valid. Ceneo simply answered a different question than the one we asked, and in a browser the same URL landed on a consent and challenge wall whose entire body text was the cookie notice.\u003C/p>\n\u003Cp>This is the failure mode that silently poisons a price dataset. If your parser is tolerant — and most are, because real pages vary — it will find zero products, report zero products, and your pipeline will interpret &quot;zero competitor offers today&quot; as a market fact rather than a fetch failure.\u003C/p>\n\u003Ch3 id=\"3-the-self-documenting-challenge\">3. The self-documenting challenge\u003C/h3>\n\u003Cp>Cdiscount was the most interesting result. \u003Ccode>200\u003C/code>, 13.7 KB, and the body was not a shop at all. It was a bot-management payload that stated its own conclusion: the request had been marked for a JavaScript challenge, and the bot category was left as \u003Ccode>unknown\u003C/code>. No offers, no structured product data, no price token. Loading the same URL in a real browser produced a page titled \u003Ccode>Accès bloqué\u003C/code>.\u003C/p>\n\u003Cp>The lesson is not that Cdiscount is hard. It is that \u003Cstrong>the block told us it was a block, in the response body, and a status-code check would have thrown that information away.\u003C/strong> The evidence you need to diagnose the failure is frequently sitting in the bytes you are not reading.\u003C/p>\n\u003Caside class=\"article__inline-cta\">\u003Cp class=\"article__inline-cta-text\">Try ScrapeWise on your own URL. \u003Cstrong>Your first 5 requests are free.\u003C/strong>\u003C/p>\u003Ca class=\"article__inline-cta-btn\" href=\"https://portal.scrapewise.ai/login\" target=\"_blank\" rel=\"noopener\">Start Free →\u003C/a>\u003C/aside>\u003Ch2 id=\"the-two-that-actually-worked\">The Two That Actually Worked\u003C/h2>\n\u003Cp>Coolblue was the surprise. A plain fetch of a product URL returned 1.4 MB carrying a structured price, currency, stock state, brand and rating, and a plain fetch of a search URL returned a structured list of 24 products with the same pricing detail on each. No browser needed for either step — discovery or reading.\u003C/p>\n\u003Cp>eMAG was similar: product tiles, prices and stock wording sitting in the raw HTML. It is one of very few large European marketplaces where that is still true, and it is worth knowing before you budget for a browser you do not need.\u003C/p>\n\u003Cp>But Coolblue also produced two traps that are precisely the kind of thing a status check will never surface — and one of them was our own bug.\u003C/p>\n\u003Cp>\u003Cstrong>Our parser lied to us first.\u003C/strong> Our initial pass reported \u003Cem>zero\u003C/em> products on Coolblue&#39;s search page, and we nearly wrote it up as a render-tier site on that basis. The response was fine. Our extraction was wrong: we were matching JSON-LD nodes with a top-level \u003Ccode>@type: &quot;Product&quot;\u003C/code>, and Coolblue nests its listing data inside an \u003Ccode>ItemList\u003C/code> under \u003Ccode>itemListElement\u003C/code>. Twenty-four products were sitting in the payload we had already downloaded. If you only check the top level of a JSON-LD graph, you will misclassify working sites as blocked ones.\u003C/p>\n\u003Cp>\u003Cstrong>The locale prefix silently discarded our query.\u003C/strong> \u003Ccode>/en/zoeken?query=philips+hue\u003C/code> returned \u003Ccode>200\u003C/code> and a full, valid page of products — Bosch vacuum cleaners. The query parameter was ignored entirely. The same query under \u003Ccode>/nl/zoeken?query=philips+hue\u003C/code> applied correctly. Status \u003Ccode>200\u003C/code>, real page, real products, completely wrong products. No status code, no content-length check and no &quot;did we get HTML&quot; assertion catches that one.\u003C/p>\n\u003Ch2 id=\"how-to-detect-a-block-without-trusting-the-status-code\">How to Detect a Block Without Trusting the Status Code\u003C/h2>\n\u003Cp>Replace your status check with content assertions. A response is a success only if it proves it contains what you came for.\u003C/p>\n\u003Col>\n\u003Cli>\u003Cstrong>Assert a currency or price token.\u003C/strong> For a product or listing page, the absence of any price-shaped string is decisive. Both Ceneo and Cdiscount fail this instantly.\u003C/li>\n\u003Cli>\u003Cstrong>Assert structured data, and walk the whole graph.\u003C/strong> Look for \u003Ccode>Product\u003C/code>, \u003Ccode>Offer\u003C/code>, \u003Ccode>ItemList\u003C/code> and \u003Ccode>itemListElement\u003C/code> — not just top-level \u003Ccode>Product\u003C/code>. This is the check that misclassified Coolblue for us.\u003C/li>\n\u003Cli>\u003Cstrong>Assert a minimum item count, and compare it to last run.\u003C/strong> A listing that returned 24 products yesterday and 0 today is a fetch failure until proven otherwise. Never let zero pass as a value.\u003C/li>\n\u003Cli>\u003Cstrong>Assert the title matches the request.\u003C/strong> Ceneo&#39;s block is only visible if you notice the page title belongs to the homepage rather than to your search. Comparing a normalised title or canonical URL against the URL you requested catches silent substitution cheaply.\u003C/li>\n\u003Cli>\u003Cstrong>Log body length and keep the first kilobyte on failure.\u003C/strong> Cdiscount&#39;s payload explained itself. You cannot read it if you threw it away.\u003C/li>\n\u003Cli>\u003Cstrong>Probe the site root separately.\u003C/strong> If the homepage also refuses you, the problem is your identity, not your URL.\u003C/li>\n\u003C/ol>\n\u003Cp>Content assertions fail loudly on real blocks and on your own parser bugs, which is the point. A status check can only ever detect the former, and it did not even manage that on half our sample.\u003C/p>\n\u003Ch2 id=\"what-it-costs-to-get-through\">What It Costs to Get Through\u003C/h2>\n\u003Cp>Six of the eight needed more than a plain request, and knowing which tier a site requires is the difference between a sane bill and a surprising one. On \u003Ca href=\"https://scrapewise.ai/pricing\">our pricing\u003C/a> a plain page is €0.15 per 1,000; a rendered page is €0.75; the render-plus-residential combination the harder sites need is €3.75.\u003C/p>\n\u003Cp>The mistake is applying the top tier everywhere. Coolblue and eMAG genuinely do not need it, and paying €3.75 per 1,000 for pages that answer a €0.15 request is a 25x overspend on the sites that are easiest to read. The opposite mistake is worse: budgeting the cheap tier for Allegro or Alza produces a pipeline that reports success and stores nothing.\u003C/p>\n\u003Cp>For Alza specifically the challenge is Cloudflare&#39;s, and the general approach is covered in our write-up on \u003Ca href=\"https://scrapewise.ai/blogs/bypass-cloudflare-akamai-perimeterx-web-scraping-2026\">bypassing Cloudflare, Akamai and PerimeterX\u003C/a>. If you are hitting \u003Ccode>__ddg\u003C/code> cookies instead, \u003Ca href=\"https://scrapewise.ai/blogs/bypass-datadome-web-scraping-2026\">DataDome behaves differently\u003C/a> and needs a different setup.\u003C/p>\n\u003Ch2 id=\"what-this-means-for-price-monitoring\">What This Means for Price Monitoring\u003C/h2>\n\u003Cp>If you run \u003Ca href=\"https://scrapewise.ai/use-cases/competitor-price-tracking\">competitor price tracking\u003C/a> across European channels, the practical consequence is that your success metric is probably wrong. Not inaccurate at the margin — structurally measuring the wrong thing on any target that prefers a quiet block to a loud one.\u003C/p>\n\u003Cp>Two changes pay for themselves immediately. Redefine a successful run as one that passed its content assertions, not one that returned \u003Ccode>200\u003C/code>. And alert on a collapse in row count per site, because that is the signal a silent substitution actually produces.\u003C/p>\n\u003Cp>Everything above came from eight requests and reading the responses. It is not sophisticated work. It is just work that a status-code check lets you skip, which is exactly why it is worth doing.\u003C/p>\n\u003Cp>If you would rather not maintain the tier logic and the assertions yourself, our \u003Ca href=\"https://scrapewise.ai/scrapers\">data APIs\u003C/a> publish real sample rows on every page so you can see the actual columns before you sign up, and the \u003Ca href=\"https://scrapewise.ai/custom-scrapers\">custom scrapers\u003C/a> cover the sites that have no ready endpoint. Every account starts with 5 free requests.\u003C/p>\n\u003Cp>\u003Cem>Nothing here is legal advice. Every request described was to a page publicly visible to any visitor, but whether you may collect and use that data depends on the site&#39;s terms, your jurisdiction and your purpose. Take your own advice before you start.\u003C/em>\u003C/p>\n","HTTP 200 Is Not Success: 8 EU Marketplaces Tested","We fetched 8 European marketplaces in September 2026. Four answered HTTP 200 and only two returned the page. What the other bodies actually contained.","ScrapeWise Team","Insights",9,[16,22,27],{"slug":17,"title":18,"image":19,"date":20,"category":13,"excerpt":21},"export-google-maps-to-excel-csv-2026","How to Export Google Maps Data to Excel & CSV: 3 Methods (and 1 to Avoid)","/img/news/export-google-maps-to-excel-csv-2026.png","15 Sep 2026","No export button? Export Google Maps to Excel or CSV with Takeout, manual copy or the Places API (60-result cap), plus EU-safe sources for bulk business data.",{"slug":23,"title":24,"image":25,"date":20,"category":13,"excerpt":26},"youtube-scraper-api-python-2026","YouTube Scraper API with Python: Official API vs Scraper, With Working Code","/img/news/youtube-scraper-api-python-2026.png","Collect YouTube channel data in Python two ways: the official API (1 unit per lookup, 10,000 free daily) or a scraper API. Working code and EU rules.",{"slug":28,"title":29,"image":30,"date":31,"category":13,"excerpt":32},"competitor-assortment-gap-analysis-2026","What Your Competitors Stock: A Practical Guide to Assortment Gap Analysis","/img/news/competitor-assortment-gap-analysis-2026.png","02 Sep 2026","Price is half the picture. A practical method for tracking which SKUs competitors carry, drop and add — and finding the range gaps before your next buying round.",{"slug":34,"title":35},"amazon-buy-box-competition-uk-2026","Amazon Buy Box Competition: 72% of Listings With Offer Data Have One Seller",[37,41,44,47,51,54,57,60,63,66],{"level":38,"text":39,"id":40},2,"Why an HTTP Status Code Cannot Tell You If a Scrape Worked","why-an-http-status-code-cannot-tell-you-if-a-scrape-worked",{"level":38,"text":42,"id":43},"What We Measured: 8 European Marketplaces, One Plain Fetch Each","what-we-measured-8-european-marketplaces-one-plain-fetch-eac",{"level":38,"text":45,"id":46},"The Three Shapes of a Block","the-three-shapes-of-a-block",{"level":48,"text":49,"id":50},3,"1. The honest refusal","1-the-honest-refusal",{"level":48,"text":52,"id":53},"2. The silent substitution","2-the-silent-substitution",{"level":48,"text":55,"id":56},"3. The self-documenting challenge","3-the-self-documenting-challenge",{"level":38,"text":58,"id":59},"The Two That Actually Worked","the-two-that-actually-worked",{"level":38,"text":61,"id":62},"How to Detect a Block Without Trusting the Status Code","how-to-detect-a-block-without-trusting-the-status-code",{"level":38,"text":64,"id":65},"What It Costs to Get Through","what-it-costs-to-get-through",{"level":38,"text":67,"id":68},"What This Means for Price Monitoring","what-this-means-for-price-monitoring",[],1790758805792]