USE CASE

Sitemap Scraper: Every Product URL From One File

A shop's sitemap is the list of pages it wants found, which makes it the cheapest way to discover a whole catalogue. Scrapewise reads the sitemap, follows the index files, filters down to product pages, and then scrapes them into one table.

THE PROBLEM

Why the Sitemap Is the Right Starting Point, and Where It Bites

  1. 01

    One sitemap is usually many files

    The standard caps a sitemap at 50,000 URLs, so large shops publish an index that points at a dozen gzipped children. A reader that only parses the first file finds a fraction of the catalogue and reports it as the whole thing.

  2. 02

    Most of it is not products

    Blog posts, category pages, filter combinations and static pages sit in the same sitemap. Scraping all of them costs pages you did not want and fills the output with rows that have no price.

  3. 03

    Sitemaps go stale in both directions

    Delisted products stay listed for weeks, and new ones appear before the sitemap is regenerated. A sitemap is a good starting list, not an inventory, so the run has to tolerate URLs that no longer resolve.

  4. 04

    Crawling the site instead is far more expensive

    Following links from the homepage means fetching every category and filter page to find the products behind them. The sitemap gives you the same destination list without paying for the pages in between.

1 URL

One sitemap address is enough to discover a whole catalogue

35

Ready data APIs for marketplaces and search, plus AI scrapers for any other shop

5

Free requests on every new account, no card needed

From a Sitemap Address to a Scraped Catalogue in 4 Steps

Discovery and scraping are two separate steps on purpose: you see the link list, and decide what it costs, before anything is scraped.

  1. Point at the sitemap

    Give the sitemap address, or just the domain and let Scrapewise find it from robots.txt. Sitemap index files are followed to their children, and gzipped sitemaps are unpacked.

  2. Filter to the pages you want

    Keep only the URLs that match a path pattern, so product pages stay and blog, category and filter URLs drop out. You see the count before the run, which is the count you will be charged for.

  3. Scrape the kept URLs

    The filtered list becomes the run's link list. Name the fields you want back: price, sale price, currency, title, stock, pack size, EAN. Pages that fail are retried and are not charged.

  4. Re-harvest on a schedule

    Re-read the sitemap daily or weekly to pick up new products and spot ones that disappeared. Group rules clean every run the same way, so the columns stay comparable over time.

What You Send, What You Get Back

You bring one address and a path filter. You get a product URL list, then a row per URL.

You give
  • Sitemap or domainrequired

    A sitemap.xml, a sitemap index, or just the domain — robots.txt is checked for the sitemap address.

    https://shop-a.example/sitemap.xml
  • Which URLs to keepoptional

    A path pattern so products stay and blog or category pages drop out. Leave it empty to keep everything.

    /product//p/-p-
  • Fields you want backrequired

    Name the columns for the scrape step. Anything the product page shows can be a column.

    pricetitlestockean
  • Re-harvest scheduleoptional
    • Daily
    • Weekly
    • On demand

    How often the sitemap is read again to pick up new and removed products.

You get

One row per kept product URL, 8 columns each

  • url
  • sitemap_file
  • lastmod
  • title
  • price
  • currency
  • stock
  • first_seen

CSV or Excel export, REST API, or MCP for AI agents.

A Sample Harvest: One Sitemap Index, 10 Files

Illustrative figures in the shape of a mid-size shop, not a customer's site. The harvest step produces the link list; the scrape step runs on the kept column.

Sitemap fileURLs listedKept after filterWhy
sitemap.xml (index)11 children—Index file, followed to its children
sitemap_products_1.xml.gz50,00048,310Product paths kept, 1,690 were redirects
sitemap_products_2.xml.gz12,74412,201Product paths kept, 543 already delisted
sitemap_categories.xml4,1200Category pages: no price to read
sitemap_filters.xml38,9000Filter combinations: duplicate products
sitemap_blog.xml8700Editorial pages
sitemap_brands.xml6120Brand landing pages
sitemap_static.xml410About, delivery, terms
sitemap_images.xml96,0040Image sitemap, not pages
Total203,29160,51160,511 product pages queued to scrape
RESULT

What You Get Instead of a Crawler

A catalogue list from one address

A catalogue list from one address

One sitemap address becomes every product URL on the site, with the index files followed and the gzip unpacked for you.

You see the count before you pay

You see the count before you pay

Harvest and scrape are separate steps, so the filtered URL count is visible first and that is the number of pages the run will cost.

New and removed products show up

New and removed products show up

Re-reading the sitemap on a schedule surfaces products that appeared and ones that stopped being listed, instead of a flat snapshot.

THE SHORT ANSWER

Sitemap Scraping: The Short Answers

How do I get all the product URLs from a website?

How do I get all the product URLs from a website?

Read its sitemap. Give Scrapewise the sitemap address or just the domain, and it follows the sitemap index to every child file, unpacks gzip, and returns the URLs. A path filter keeps the product pages and drops blog and category URLs.

What if the sitemap is split into several files?

What if the sitemap is split into several files?

That is the normal case above 50,000 URLs, and it is handled. The top-level file is a sitemap index listing children, often gzipped. All of them are followed and merged into one link list, with the source file kept as a column.

Is reading a sitemap allowed?

Is reading a sitemap allowed?

A sitemap is published by the site for crawlers to read and its address is usually declared in robots.txt. Scrapewise reads public pages only, at a rate that does not interfere with the site, and does not touch anything behind a login.

Give Us One Sitemap Address

Point Scrapewise at a sitemap, see the filtered product URL count, then decide what to scrape. Every new account gets 5 free requests, and after that you pay as you go from a prepaid balance. Failed pages are not charged.

FAQ

Sitemap Scraper FAQs

How the sitemap is read, what gets filtered, and what it costs.

A sitemap scraper reads a site's sitemap.xml to discover the pages it publishes, then scrapes those pages for data. It is the cheapest way to cover a whole catalogue, because the site has already listed every product URL for crawlers to find.