[{"data":1,"prerenderedAt":85},["ShallowReactive",2],{"$f-N7yc8RUZouj7Kl4UysD627kORgpBWtiysCvsC7Z71k":3},{"title":4,"date":5,"dateModified":6,"datePublished":7,"dateModifiedISO":7,"image":8,"content":9,"faq":10,"metaTitle":30,"metaDescription":31,"author":32,"authorBio":33,"authorLinkedin":33,"authorTitle":33,"authorPhoto":33,"lastReviewed":33,"researchBasis":33,"category":34,"readingTime":35,"related":36,"prev":52,"next":55,"toc":58,"takeaways":84},"LLM Data Extraction: Let the AI Read, Let Rules Do the Maths","30 Sep 2026","30 SEP 2026","2026-09-30","/img/news/llm-data-extraction-rules-do-the-maths-2026.png","\u003Cp>We asked an LLM extraction step to read a yarn shop&#39;s product page and return the price per 50 g. It returned a number on almost every row. The number was right on 18 of 69 rows.\u003C/p>\n\u003Cp>That sounds like a model that is right a quarter of the time. It is worse than that. When we checked every row, not one of the 18 held a real calculation. They were all 50 g balls, where the price per 50 g is the price. The model had copied the price, and on those rows copying happened to be correct.\u003C/p>\n\u003Cp>This post is about LLM data extraction accuracy on the one thing it is bad at, and the split we use now: the AI reads raw values off the page, and plain rules do every calculation.\u003C/p>\n\u003Ch2 id=\"the-setup\">The Setup\u003C/h2>\n\u003Cp>The job was a price per unit comparison for a yarn retailer, using the competitor links they gave us. The full story is in \u003Ca href=\"/blogs/price-per-unit-comparison-pack-sizes-2026\">comparing prices across pack sizes\u003C/a>. The schema for the AI scraper asked for name, price, sale price, currency, pack weight in grams, price per 50 g and sale price per 50 g.\u003C/p>\n\u003Cp>The first run looked fine at a glance. Every column was filled. Most values were plausible. You would only see the problem by recomputing each row, which is exactly the check nobody does on a busy day.\u003C/p>\n\u003Caside class=\"article__usecase-card\">\u003Cdiv class=\"article__usecase-label\">Related use case\u003C/div>\u003Ch3 class=\"article__usecase-title\">Any-site data scraper\u003C/h3>\u003Cp class=\"article__usecase-blurb\">No-code extraction from any website. Managed infrastructure, no anti-bot headaches.\u003C/p>\u003Ca class=\"article__usecase-link\" href=\"/use-cases/data-scraper\">See how it works →\u003C/a>\u003C/aside>\u003Ch2 id=\"what-the-model-wrote-in-the-price-per-50-g-column\">What the Model Wrote in the Price per 50 g Column\u003C/h2>\n\u003Cp>We went through all 69 yarn rows by hand against the live pages.\u003C/p>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>What the column held\u003C/th>\n\u003Cth>Rows\u003C/th>\n\u003C/tr>\n\u003C/thead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>The price, copied. Right only because the pack is 50 g\u003C/td>\n\u003Ctd>18\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>The price, copied, although the pack is not 50 g\u003C/td>\n\u003Ctd>22\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>About price ÷ 5, a formula nobody asked for\u003C/td>\n\u003Ctd>11\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>The other printed price on the page (the before-price or the sale price)\u003C/td>\n\u003Ctd>6\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>The shop&#39;s own printed price per kg\u003C/td>\n\u003Ctd>3\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Other numbers from the page\u003C/td>\n\u003Ctd>3\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>Empty, because price or pack was missing\u003C/td>\n\u003Ctd>6\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>A real, correct calculation\u003C/td>\n\u003Ctd>0\u003C/td>\n\u003C/tr>\n\u003C/tbody>\u003C/table>\n\u003Cp>The sale price per 50 g was worse. On 49 rows with no sale at all, the model filled it with a copy of the price. So on a chart, most products looked like they were on sale at their normal price.\u003C/p>\n\u003Cp>We rewrote the schema twice. Once with the formula and an example in the field description, once with an instruction at page level to calculate. Neither changed the pattern. The model copies numbers it can see. It does not produce a number that is not on the page.\u003C/p>\n\u003Ch2 id=\"the-pack-weight-was-also-wrong-for-different-reasons\">The Pack Weight Was Also Wrong, for Different Reasons\u003C/h2>\n\u003Cp>The per-50 g value needs a correct pack weight, and that was wrong on 25 of 67 links. These errors were not arithmetic. They were reading errors, and they had causes you can recognise:\u003C/p>\n\u003Cul>\n\u003Cli>\u003Cstrong>Another number on the page.\u003C/strong> A free-shipping threshold of 699 became the pack weight. A pattern note, &quot;about 450 g for a sweater&quot;, became 450.\u003C/li>\n\u003Cli>\u003Cstrong>Several weights on one page.\u003C/strong> A product page listing several colourways with different ball sizes gave the weight of a colourway that was not the linked one. That happened on 6 links.\u003C/li>\n\u003Cli>\u003Cstrong>Units.\u003C/strong> Shops that store weight in kilograms gave 0.05 and 0.1, and the model wrote those as grams.\u003C/li>\n\u003Cli>\u003Cstrong>Nothing found.\u003C/strong> On 4 pages the price and the weight were in the page, but not where the model looked, for example inside microdata or written by a script.\u003C/li>\n\u003Cli>\u003Cstrong>The wrong product.\u003C/strong> A collection page with many yarns gave a different yarn.\u003C/li>\n\u003C/ul>\n\u003Cp>One of the 25 was not the model&#39;s fault at all: the customer&#39;s file was out of date, and the shop&#39;s 226 g was right.\u003C/p>\n\u003Caside class=\"article__inline-cta\">\u003Cp class=\"article__inline-cta-text\">Try ScrapeWise on your own URL. \u003Cstrong>Your first 5 requests are free.\u003C/strong>\u003C/p>\u003Ca class=\"article__inline-cta-btn\" href=\"https://portal.scrapewise.ai/login\" target=\"_blank\" rel=\"noopener\">Start Free →\u003C/a>\u003C/aside>\u003Ch2 id=\"why-this-happens\">Why This Happens\u003C/h2>\n\u003Cp>An extraction model is trained to find a value in text and return it in a field. That is reading. A price per unit is a derived value: it is not printed on the page, so there is nothing to find, and the model fills the field with the closest thing it can find. The output format hides the difference. A copied number and a calculated number look the same in a JSON field.\u003C/p>\n\u003Cp>You can see the same thing in other LLM work. Models are good at pulling &quot;2.60&quot; and &quot;100 g&quot; out of messy text. They are unreliable at doing the division in the same step, and they give no signal when they skip it.\u003C/p>\n\u003Ch2 id=\"the-split-that-works\">The Split That Works\u003C/h2>\n\u003Cp>We changed the rule for every scraper. The AI copies what is printed, with no maths and no unit changes. Every calculation is an after-scrape rule, which is plain code that does the same thing on every row.\u003C/p>\n\u003Cp>The raw schema now has: name, pack_grams, price, sale_price, currency, availability. Every field description says &quot;copy as printed&quot;.\u003C/p>\n\u003Cp>The rules then do the rest:\u003C/p>\n\u003Col>\n\u003Cli>\u003Cstrong>Take a number from text\u003C/strong> reads the number in front of a unit word, such as &quot;Wool 100 g&quot; in a product name, and can convert kg to g. We use it where the weight is only in text.\u003C/li>\n\u003Cli>\u003Cstrong>Price per unit\u003C/strong> computes \u003Ccode>price ÷ pack_grams × 50\u003C/code>. It skips the row when the price or the pack is missing, so an unknown stays empty instead of becoming 0.\u003C/li>\n\u003Cli>\u003Cstrong>Convert currency\u003C/strong> turns the result into EUR with the ECB rate of the day, and writes the rate and its date next to it.\u003C/li>\n\u003Cli>\u003Cstrong>Replace values from a list\u003C/strong> maps each shop&#39;s stock wording to IN_STOCK, OUT_OF_STOCK, BACKORDER or UNKNOWN.\u003C/li>\n\u003C/ol>\n\u003Cp>After the change the arithmetic was right on every row that had a price and a pack, 100% in the first run and 100% in the final setup.\u003C/p>\n\u003Ch2 id=\"what-a-text-cleaning-rule-cannot-do\">What a Text-Cleaning Rule Cannot Do\u003C/h2>\n\u003Cp>Before the &quot;Price per unit&quot; and &quot;Take a number from text&quot; rules existed, we tried to get by with the text-cleaning rule we had, which keeps only the characters you choose. It is useful for codes. It is dangerous for numbers. We tested it on real pack texts:\u003C/p>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Input\u003C/th>\n\u003Cth>Keep\u003C/th>\n\u003Cth>Output\u003C/th>\n\u003Cth>Right?\u003C/th>\n\u003C/tr>\n\u003C/thead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>\u003Ccode>100 g\u003C/code>\u003C/td>\n\u003Ctd>digits\u003C/td>\n\u003Ctd>\u003Ccode>100\u003C/code>\u003C/td>\n\u003Ctd>Yes\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>0.05 kg\u003C/code>\u003C/td>\n\u003Ctd>digits, dot, comma\u003C/td>\n\u003Ctd>\u003Ccode>0.05\u003C/code>\u003C/td>\n\u003Ctd>The digits, but still kg\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>100m/50g\u003C/code>\u003C/td>\n\u003Ctd>digits\u003C/td>\n\u003Ctd>\u003Ccode>10050\u003C/code>\u003C/td>\n\u003Ctd>No, metres and grams glued together\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>3.53 oz\u003C/code>\u003C/td>\n\u003Ctd>digits\u003C/td>\n\u003Ctd>\u003Ccode>353\u003C/code>\u003C/td>\n\u003Ctd>No, the decimal point is gone\u003C/td>\n\u003C/tr>\n\u003Ctr>\n\u003Ctd>\u003Ccode>49,95 kr.\u003C/code>\u003C/td>\n\u003Ctd>digits, dot, comma\u003C/td>\n\u003Ctd>\u003Ccode>49,95.\u003C/code>\u003C/td>\n\u003Ctd>No, a trailing dot and a comma decimal\u003C/td>\n\u003C/tr>\n\u003C/tbody>\u003C/table>\n\u003Cp>This is why the number rule reads the number in front of a named unit and ignores the other digits on the line.\u003C/p>\n\u003Ch2 id=\"use-less-ai-where-a-structured-source-exists\">Use Less AI Where a Structured Source Exists\u003C/h2>\n\u003Cp>The second fix was to use the AI less. Many of the shops had a structured source with the weight in grams already as a number: Shopify&#39;s variant JSON, WooCommerce&#39;s Store API, or \u003Ccode>itemprop\u003C/code> microdata in the HTML. For those links we replaced the AI scraper with an API or microdata scraper. Under 10% of the links that return data still use AI, because no structured source on those shops has the weight.\u003C/p>\n\u003Cp>We choose the source in this order: long-run reliability first, then scale, then cost. Structured sources win on all three. On several live pages the AI step failed and returned no data, while a scraper without the AI step read the same pages at the cheapest fetch tier. And on the first Shopify links we moved, the JSON run cost about a ninth of the AI run. More on finding these sources in \u003Ca href=\"/blogs/find-hidden-json-api-shop-page-2026\">how to find the hidden JSON API behind a shop page\u003C/a>.\u003C/p>\n\u003Ch2 id=\"a-checklist-for-ai-extraction\">A Checklist for AI Extraction\u003C/h2>\n\u003Cp>If you use LLM extraction for prices, this is what we would do from day one:\u003C/p>\n\u003Cul>\n\u003Cli>Ask only for values that are printed on the page. No totals, ratios, conversions or &quot;per&quot; values.\u003C/li>\n\u003Cli>Put every calculation in a rule you can read, and test the rule on sample values before you save it.\u003C/li>\n\u003Cli>Leave a derived value empty when an input is missing. Never 0.\u003C/li>\n\u003Cli>Recompute a sample of rows by hand after the first run. A filled column is not a correct column.\u003C/li>\n\u003Cli>Check that the sale price is lower than the regular price, and treat &quot;equal&quot; as no sale.\u003C/li>\n\u003Cli>If a structured source exists, prefer it over AI for the fields it has.\u003C/li>\n\u003C/ul>\n\u003Cp>AI extraction is still the right tool for pages with no structure at all, and it saved us a lot of selector writing on the links where nothing else worked. We just stopped asking it to do sums. If you want to try the same split in your own account, the \u003Ca href=\"/use-cases/data-scraper\">data scraper\u003C/a> takes a URL and suggests the fields, and the \u003Ca href=\"/blogs/build-web-scraper-with-claude-mcp-2026\">MCP walkthrough\u003C/a> shows how to have Claude set up the rules for you. For the general trade-offs of AI in scraping, see \u003Ca href=\"/blogs/ai-powered-web-scraping-2026\">AI-powered web scraping\u003C/a>.\u003C/p>\n",{"title":11,"description":12,"badge":13,"benefits":14},"Frequently asked questions","LLM data extraction accuracy: questions answered","FAQ",[15,18,21,24,27],{"title":16,"description":17},"Why did the LLM get the price per unit wrong?","An extraction model copies values it can see on the page. A price per unit is not printed, so the model filled the field with the nearest number, usually the price itself. It never did the division.",{"title":19,"description":20},"What should an AI extraction step return?","Only values that are printed on the page, copied as they appear: name, price, sale price, currency, pack size and stock wording. Calculations and unit changes belong in rules.",{"title":22,"description":23},"What are after-scrape rules?","They are fixed steps that run on every scraped row, such as converting currency, mapping stock wording, cleaning codes, taking a number from text and rescaling a price to a fixed pack size.",{"title":25,"description":26},"How accurate was the AI on pack weight?","It was wrong on 25 of 67 links. It read other numbers such as a free shipping threshold, took the weight of another colourway, or wrote kilograms as grams.",{"title":28,"description":29},"When is AI extraction still the right choice?","When a page has no API, no JSON-LD and no microdata. In our yarn project that was under 10% of the links, and there the AI copies raw values only.","LLM Data Extraction: Let AI Read, Let Rules Do the Maths","Our AI extraction got a price per 50 g right on 18 of 69 rows, and never by calculating. Moving the maths into rules fixed it on every row.","Raivo Kartau",null,"Scraping",7,[37,42,47],{"slug":38,"title":39,"image":40,"date":5,"category":34,"excerpt":41},"build-web-scraper-with-claude-mcp-2026","Build a Web Scraper with Claude and MCP: A Step-by-Step Walkthrough","/img/news/build-web-scraper-with-claude-mcp-2026.png","Connect Claude to Scrapewise through MCP and let it build, test and run price scrapers for you. The setup, the prompts and the guardrails we use.",{"slug":43,"title":44,"image":45,"date":5,"category":34,"excerpt":46},"find-hidden-json-api-shop-page-2026","How to Find the Hidden JSON API Behind a Shop Page (and Why It Is Cheaper)","/img/news/find-hidden-json-api-shop-page-2026.png","Most shop pages load prices from JSON first. How we find those endpoints (Shopify, WooCommerce, Magento, JSON-LD, search APIs) and test them.",{"slug":48,"title":49,"image":50,"date":5,"category":34,"excerpt":51},"web-scraping-mistakes-price-scrapers-2026","14 Web Scraping Mistakes We Made Building Price Scrapers","/img/news/web-scraping-mistakes-price-scrapers-2026.png","Wrong sources, shifted EANs, Excel eating a brand name, AI doing sums. 14 mistakes from two real price monitoring builds, and the fix for each.",{"slug":53,"title":54},"price-per-unit-comparison-pack-sizes-2026","Price per Unit: How to Compare Competitor Prices Across Pack Sizes",{"slug":56,"title":57},"http-200-not-success-eu-marketplaces-2026","HTTP 200 Doesn't Mean You Got the Page: 8 European Marketplaces Measured",[59,63,66,69,72,75,78,81],{"level":60,"text":61,"id":62},2,"The Setup","the-setup",{"level":60,"text":64,"id":65},"What the Model Wrote in the Price per 50 g Column","what-the-model-wrote-in-the-price-per-50-g-column",{"level":60,"text":67,"id":68},"The Pack Weight Was Also Wrong, for Different Reasons","the-pack-weight-was-also-wrong-for-different-reasons",{"level":60,"text":70,"id":71},"Why This Happens","why-this-happens",{"level":60,"text":73,"id":74},"The Split That Works","the-split-that-works",{"level":60,"text":76,"id":77},"What a Text-Cleaning Rule Cannot Do","what-a-text-cleaning-rule-cannot-do",{"level":60,"text":79,"id":80},"Use Less AI Where a Structured Source Exists","use-less-ai-where-a-structured-source-exists",{"level":60,"text":82,"id":83},"A Checklist for AI Extraction","a-checklist-for-ai-extraction",[],1790769637106]