REWE Scraper
We do not have a dedicated REWE data API. What we do have is an AI scraper that reads a shop.rewe.de page you point it at, turns it into the columns you asked for and hands them back by REST, CSV or Excel — on a schedule, priced per page.
- shop.rewe.de
How to scrape REWE
REWE publishes no read API for its catalogue. There is no key to apply for and no endpoint to call, which is why the page is the interface. REWE is the German grocery price reference and prices at store level, which makes a national German grocery number a fiction. The way to get the data is to read the page the way a customer sees it: point the AI scraper at a shop.rewe.de URL, declare the columns you want in plain English, and schedule it. Prices here are set per store, so a row is only meaningful with the store it came from attached, and a scrape that ignores store selection is quietly averaging several different markets.
What the AI scraper does with a shop.rewe.de link
The same three steps as any other page you point it at. Nothing about REWE is pre-built, which is exactly why it works on pages nobody wrote a connector for.
Paste the link
Give it a shop.rewe.de product or listing URL. Market, currency and assortment all come from the page itself, so there is no region setting to get wrong — the URL decides.
Declare the columns
Name each field, give it a type and write a line telling the AI what to look for. “product_title” and “brand” are the two most people start with. The schema is yours and you can change it between runs.
Schedule it and collect
Run it hourly, daily or on demand. Every run returns the same columns in the same order with a timestamp on each row, so the history builds itself.
Set it up once. The data keeps coming.
Save your search once and it runs by itself. Every run lands in your own Scrapewise database, and you pick how to read it.
Set it up once
Add your ASINs, keywords, places or apps to a scraper in the web app. New accounts get 5 free requests.
Runs on your schedule
Pick daily, weekly on the days you choose, every few days, or the first or last day of the month. Scheduled runs start at night, European time. Need it more often? Start runs from your own code.
Saved in your database
Each run adds dated rows, so this week sits next to last week. Rows are kept for 90 days.
Use it your way
Open the table in the web app and download the latest run as an Excel file. Ask your own AI assistant about it. Or pull the rows into your code with an API key.
Ask your AI assistant about your own data
Connect Claude Desktop or Claude Code with a read-only key. The assistant reads the rows you've already collected, at no extra cost on Scrapewise. With a full-access key it can start a run for you too.
- Which of my ASINs lost the Buy Box this week?
- Which competitor cut prices the most since Monday?
- Show the keywords where I dropped out of the top 10.
What you send and what you get
You send a link and a column list. Everything else — market, currency, which products are on the page — comes from the page, so there is nothing to configure twice.
- Link to a REWE pagerequired
Open the product or category page in your browser and copy the address. Any REWE country domain works the same way.
https://www.shop.rewe.de/ - Your column listrequired
Field name, type and one line of plain English per column. Written once, reused on every run.
product_title, brand, store_id, pack_size - How often it runsoptional
Hourly, daily, weekly or on demand through the API. Daily is what most price monitoring uses.
Daily at 06:00 UTC
One row per product for the selected store, 12 columns each
- product_title
- brand
- store_id
- pack_size
- price
- unit_price
- unit
- loyalty_price
- currency
- availability
- product_url
- checked_at
REST API, CSV or Excel. The column set stays the same between runs, so downstream jobs do not re-map fields.
The columns a REWE run comes back with
We are not going to print a table of invented grocery (store-level pricing) prices and call it real output. REWE answers an ordinary request with Cloudflare rather than a product page, so we have no run of our own to quote from. So what is published here is the column list you declare before the run, which is the same list every run comes back with.
product_titleProduct title as the retailer writes itbrandBrand as the retailer labels itstore_idWhich store or ZIP this price was quoted forpack_sizePack size as printedpriceShelf price for that store todayunit_pricePrice per unit, where the page prints itunitWhich unit that per-unit price is quoted inloyalty_priceMember or card price, where the page shows onecurrencyCurrency the page quotedavailabilityWhether the line is in stock at that storeproduct_urlLink back to the page the row came fromchecked_atWhen this run collected the page (UTC)Every column here is one you name yourself in the schema. If a field is not on the page on the day of the run, it comes back empty rather than guessed.
What a REWE page costs
Pay-as-you-go from your wallet. No plan, no monthly fee, no seat count.
Measured on 3 October 2026, not assumed. A plain request for a shop.rewe.de page is refused with HTTP 403 by Cloudflare, and so is the bare homepage, which rules out a bad path as the explanation. Clearing that needs both a real browser and a residential exit.
500 REWE pages checked once a day for a month is 15,000 pages.
about EUR 56.25 for the month
You are charged for what a run actually used, not what it was expected to need, and your balance never expires.
See the full price list →What this page is not promising
We would rather say this here than in a support ticket.
No dedicated REWE endpoint
There is no REWE API in our catalogue and no pre-built REWE schema. The AI scraper reads the page you point it at — that is the whole mechanism, and it is why it works on pages nobody built a connector for.
Nothing behind a login
It reads what a visitor can see. Account pricing, contract pricing and anything behind a sign-in are out of scope.
No field the page does not show
If the page does not print it, the scraper cannot return it. Identifiers like EAN or GTIN only come back where the retailer publishes them.
A price is a point in time
Every row carries the timestamp of the run that collected it. A price without one is not evidence, which is why the column is in the default schema.
REWE has one specific trap
No price renders until a Markt is selected, and the selection is held in a cookie rather than the URL.
Not a bulk dump of the catalogue
You give it the URLs you care about. It does not crawl shop.rewe.de end to end, and a schedule that tried to would cost more than the answer is worth.
Typical fields the AI scraper extracts from a REWE page
Treat this as a starting point rather than a fixed schema. What comes back is whatever the page actually shows on the day of the run.
Store is part of the price
The field most grocery scrapes omit and then cannot explain their own numbers without.
- Which store or ZIP the price was quoted for
- Shelf price for that store
- Member or loyalty price where shown
- Whether the line is in stock at that store
- Pickup or delivery availability, where the page shows it
Unit price, not shelf price
Still the only number that compares across brands, and still printed in small type.
- Pack size exactly as printed
- Price per unit where the page prints it
- Which unit the per-unit price is quoted in
- Multibuy or promotional price where shown
- Any was-price the page prints
Product identity
Own-label lines carry no manufacturer identifier, so the retailer's own code is the join key.
- Retailer's own product code
- EAN or GTIN, where published
- Brand, or own-label name
- Category breadcrumb
- Product image URLs
The column list is one you write: name each field, give it a type and a line telling the AI what to look for. The schema belongs to your run, not to us. How Custom Schema works →
What it cannot give you If a field is not visible on the page, the scraper cannot invent it. Identifiers such as EAN or GTIN only come back when REWE publishes them.
What people use REWE data for
Store-level price mapping
The same line at two of this retailer's own stores is frequently two different prices. A store column turns that from an anecdote into a map, and it is the input regional pricing actually needs.
Unit-price comparison across retailers
One schema with pack size and per-unit price across several grocers turns shelf prices into a table that genuinely compares.
Promotion calendar reconstruction
When a promotion starts, how deep it goes, how long it runs, and whether it runs everywhere or only in some stores.
Loyalty-price visibility
Member prices are a second price list sitting next to the first one. Capturing both is the difference between modelling the shelf and modelling what shoppers pay.
Availability and delisting alerts
Suppliers usually find out a line has been delisted weeks late. A daily check against a store list tells you in a day.
Feeding your own pricing rules
The export is a normal REST API or CSV, so store, line and per-unit price land directly in whatever pricing model you already run.
What makes REWE harder than an ordinary storefront
Here is what the scraper is actually up against on shop.rewe.de, measured rather than assumed.
Access
- Plain HTTP requests are refused at the edge — a naive fetch returns HTTP 403, not a page
- Named protection on the response: Cloudflare
- No usable sitemap directive in robots.txt, so the URL list comes from your own category pages
There is no single price
- Price is set per store, so the store or ZIP has to be part of the input and part of every row
- Store selection is usually held in a cookie or a session, so a run has to be pointed at a store-scoped URL rather than a generic one
- Covering several stores multiplies the page count, which is what drives the cost
- Member and loyalty prices sit alongside shelf prices and are a separate column, not a replacement for one
What the page publishes
- No product-level structured data is published, so every field is read off the rendered page
- No price renders until a Markt is selected, and the selection is held in a cookie rather than the URL.
- Price is one of the last things the page settles on, so a read taken too early records the placeholder rather than the number
Keeping a history
- A single read is a snapshot; the value is in the series, which means the run has to be scheduled and the rows kept
- Every row carries the timestamp of the run that produced it, so two days can be compared without guesswork
- Columns stay stable between runs, so a dashboard written once does not break when the site redesigns
REWE scraping — questions
What people ask before pointing the scraper at shop.rewe.de.
REWE publishes no read API for its catalogue. There is no key to apply for and no endpoint to call, which is why the page is the interface.
Ready to pull REWE data into your stack?
Start free — or talk to our team about your exact fields, refresh cadence and volume.