Custom scraper

Lidl Scraper

We do not have a dedicated Lidl data API. What we do have is an AI scraper that reads a lidl.de page you point it at, turns it into the columns you asked for and hands them back by REST, CSV or Excel — on a schedule, priced per page.

  • lidl.de
  • lidl.co.uk
  • lidl.com
Three stages of a Lidl run: paste a lidl.de URL billed on the basic tier at EUR 0.50 per 1,000 pages; declare columns such as product_title and brand; and the 11 columns it returns.

How to scrape Lidl

Lidl publishes no read API for its catalogue. There is no key to apply for and no endpoint to call, which is why the page is the interface. Lidl is the price floor in most of the European markets it operates in, which makes it the single most-watched grocery price list on the continent. The way to get the data is to read the page the way a customer sees it: point the AI scraper at a lidl.de URL, declare the columns you want in plain English, and schedule it. Grocery is the category where the shelf price tells you least: the comparable number is the price per unit, and it is frequently printed in smaller type next to a pack size that changes.

What the AI scraper does with a lidl.de link

The same three steps as any other page you point it at. Nothing about Lidl is pre-built, which is exactly why it works on pages nobody wrote a connector for.

  1. Paste the link

    Give it a lidl.de product or listing URL. Market, currency and assortment all come from the page itself, so there is no region setting to get wrong — the URL decides.

  2. Declare the columns

    Name each field, give it a type and write a line telling the AI what to look for. “product_title” and “brand” are the two most people start with. The schema is yours and you can change it between runs.

  3. Schedule it and collect

    Run it hourly, daily or on demand. Every run returns the same columns in the same order with a timestamp on each row, so the history builds itself.

Set it up once. The data keeps coming.

Save your search once and it runs by itself. Every run lands in your own Scrapewise database, and you pick how to read it.

  1. Set it up once

    Add your ASINs, keywords, places or apps to a scraper in the web app. New accounts get 5 free requests.

  2. Runs on your schedule

    Pick daily, weekly on the days you choose, every few days, or the first or last day of the month. Scheduled runs start at night, European time. Need it more often? Start runs from your own code.

  3. Saved in your database

    Each run adds dated rows, so this week sits next to last week. Rows are kept for 90 days.

  4. Use it your way

    Open the table in the web app and download the latest run as an Excel file. Ask your own AI assistant about it. Or pull the rows into your code with an API key.

Ask your AI assistant about your own data

Connect Claude Desktop or Claude Code with a read-only key. The assistant reads the rows you've already collected, at no extra cost on Scrapewise. With a full-access key it can start a run for you too.

  • Which of my ASINs lost the Buy Box this week?
  • Which competitor cut prices the most since Monday?
  • Show the keywords where I dropped out of the top 10.

What you send and what you get

You send a link and a column list. Everything else — market, currency, which products are on the page — comes from the page, so there is nothing to configure twice.

You give
  • Link to a Lidl pagerequired

    Open the product or category page in your browser and copy the address. Any Lidl country domain works the same way.

    https://www.lidl.de/
  • Your column listrequired

    Field name, type and one line of plain English per column. Written once, reused on every run.

    product_title, brand, pack_size, price
  • How often it runsoptional

    Hourly, daily, weekly or on demand through the API. Daily is what most price monitoring uses.

    Daily at 06:00 UTC
You get

One row per product on the page, 7 columns each

  • product_title
  • price
  • currency
  • availability
  • brand
  • sku
  • source_url

REST API, CSV or Excel. The column set stays the same between runs, so downstream jobs do not re-map fields.

What a run on Lidl actually returned

This is real output, not a mock-up. 4 rows, collected on 3 October 2026 from live lidl.de product pages through our own engine, the same way one of your runs would. Prices move, so treat the numbers as a sample of shape rather than a current price list.

All 7 columns

  • product_title
  • price
  • currency
  • availability
  • brand
  • sku
  • source_url

4 sample rows, all 7 columns seen across sample calls.

4 sample rows, all 7 columns seen across sample calls.
product_titleproduct_titleProduct title as the retailer writes itpricepricePrice the page is charging todaycurrencycurrencyCurrency the page quotedavailabilityavailabilityStock state the page publishedbrandbrandBrand as the retailer labels itskuskuThe retailer's own product codesource_urlsource_urlThe exact page this row was read from
Damen Trenchjacke26.99EUROnlineOnlyESMARA®100408708https://www.lidl.de/p/esmara-damen-trenchjacke/p100408708
Nähmaschine SNM 33 C1119.99EUROnlineOnly100392615https://www.lidl.de/p/naehmaschine-snm-33-c1/p100392615
Ouzo 12 38% Vol11.99EUROnlineOnly100260303https://www.lidl.de/p/ouzo-12-38-vol/p100260303
Fitnessausrüstung5.99EUROnlineOnlyCRIVIT100384035003https://www.lidl.de/p/crivit-fitnessausruestung/p100384035003

Column names are ones we chose when declaring the schema for this run. Yours can be named whatever your downstream job already expects. We pointed the run at 14 Lidl product pages and 4 came back with a readable price. The rest returned something other than a product page, which is ordinary on this storefront and the reason the tier named above is what it is. We publish what we measured, not what we expect.

What a Lidl page costs

Pay-as-you-go from your wallet. No plan, no monthly fee, no seat count.

per 1,000 pagespay-as-you-go, no plan

Measured on 3 October 2026, not assumed. A plain HTTP request for a lidl.de page comes back with the page, so most runs never need a browser and never need a residential exit.

Worked example

500 Lidl pages checked once a day for a month is 15,000 pages.

about EUR 7.50 for the month

You are charged for what a run actually used, not what it was expected to need, and your balance never expires.

See the full price list →

What this page is not promising

We would rather say this here than in a support ticket.

  • No dedicated Lidl endpoint

    There is no Lidl API in our catalogue and no pre-built Lidl schema. The AI scraper reads the page you point it at — that is the whole mechanism, and it is why it works on pages nobody built a connector for.

  • Nothing behind a login

    It reads what a visitor can see. Account pricing, contract pricing and anything behind a sign-in are out of scope.

  • No field the page does not show

    If the page does not print it, the scraper cannot return it. Identifiers like EAN or GTIN only come back where the retailer publishes them.

  • A price is a point in time

    Every row carries the timestamp of the run that collected it. A price without one is not evidence, which is why the column is in the default schema.

  • Lidl has one specific trap

    Each country has its own domain, catalogue and price, and the weekly non-food range disappears from the site once it sells through.

  • Not a bulk dump of the catalogue

    You give it the URLs you care about. It does not crawl lidl.de end to end, and a schedule that tried to would cost more than the answer is worth.

Typical fields the AI scraper extracts from a Lidl page

Treat this as a starting point rather than a fixed schema. What comes back is whatever the page actually shows on the day of the run.

Unit price, not shelf price

The only grocery number that compares across brands. A 750 ml bottle at 4.50 and a 1 litre bottle at 5.40 is not a comparison until you divide.

  • Shelf price
  • Pack size exactly as printed
  • Price per litre, kilo or unit where the page prints it
  • Which unit the per-unit price is quoted in
  • Multibuy, loyalty or promotional price where shown

Promotion mechanics

Grocery promotions are rarely a simple percentage, and the mechanic is what moves volume.

  • Promotional price and the wording that explains it
  • Multibuy conditions, where the page states them
  • Loyalty-card price where the page shows one for members
  • Promotion end date, where printed
  • Any was-price

Product identity

Own-label lines have no manufacturer identifier at all, which is why the retailer's own code matters here.

  • Retailer's own product code
  • EAN or GTIN, where published
  • Brand, or own-label name
  • Category breadcrumb
  • Product image URLs

The column list is one you write: name each field, give it a type and a line telling the AI what to look for. The schema belongs to your run, not to us. How Custom Schema works →

What it cannot give you If a field is not visible on the page, the scraper cannot invent it. Identifiers such as EAN or GTIN only come back when Lidl publishes them.

What people use Lidl data for

01

Unit-price comparison across retailers

The only honest grocery price comparison. One schema with pack size and per-unit price across several retailers turns shelf prices into a table that actually compares.

02

Promotion calendar reconstruction

When a competitor starts a promotion on a line, how deep it goes and how long it runs. That calendar is public but only visible if you collect it daily.

03

Own-label versus brand price gap

The gap between a retailer's own label and the branded equivalent is a core grocery pricing input, and it moves. This is how you measure it rather than estimate it.

04

Availability and range checks

Whether your line is actually listed, in stock and findable at a given retailer. Suppliers usually discover delisting late; a daily check tells you in a day.

05

Price-index reporting

A fixed basket of lines, read daily, is a price index you own, rather than one you buy from a panel provider with a six-week lag.

06

Feeding your own pricing rules

The export is a normal REST API or CSV, so a per-unit price lands directly in whatever pricing model you already run.

What makes Lidl harder than an ordinary storefront

Here is what the scraper is actually up against on lidl.de, measured rather than assumed.

Access

  • Plain HTTP requests are answered, but the answer is the shell of the page rather than the priced version of it
  • No named protection vendor on the response, which is why this one lands on a cheaper tier than most of the cohort
  • robots.txt publishes sitemaps, which is the cheapest way to build the URL list a run works through

Pack size is the whole problem

  • Shelf price is not comparable across brands until it is divided by a pack size the page prints in small type
  • Pack sizes change without the product changing, which silently breaks a per-unit comparison built on a fixed assumption
  • Promotions are multibuys and loyalty prices rather than simple percentages, so the mechanic has to be captured as text
  • Grocery catalogues run to tens of thousands of lines, so scope the page count before switching a run on

What the page publishes

  • Structured data is published but not at product level (MemberProgram, MemberProgramTier, Organization), so price comes from the rendered page rather than a feed
  • Each country has its own domain, catalogue and price, and the weekly non-food range disappears from the site once it sells through.
  • Price is one of the last things the page settles on, so a read taken too early records the placeholder rather than the number

Keeping a history

  • A single read is a snapshot; the value is in the series, which means the run has to be scheduled and the rows kept
  • Every row carries the timestamp of the run that produced it, so two days can be compared without guesswork
  • Columns stay stable between runs, so a dashboard written once does not break when the site redesigns
FAQ

Lidl scraping — questions

What people ask before pointing the scraper at lidl.de.

Lidl publishes no read API for its catalogue. There is no key to apply for and no endpoint to call, which is why the page is the interface.

Ready to pull Lidl data into your stack?

Start free — or talk to our team about your exact fields, refresh cadence and volume.