LLM Data Extraction: Let the AI Read, Let Rules Do the Maths

Last updated: 30 SEP 2026

LLM Data Extraction: Let the AI Read, Let Rules Do the Maths

We asked an LLM extraction step to read a yarn shop's product page and return the price per 50 g. It returned a number on almost every row. The number was right on 18 of 69 rows.

That sounds like a model that is right a quarter of the time. It is worse than that. When we checked every row, not one of the 18 held a real calculation. They were all 50 g balls, where the price per 50 g is the price. The model had copied the price, and on those rows copying happened to be correct.

This post is about LLM data extraction accuracy on the one thing it is bad at, and the split we use now: the AI reads raw values off the page, and plain rules do every calculation.

The Setup

The job was a price per unit comparison for a yarn retailer, using the competitor links they gave us. The full story is in comparing prices across pack sizes. The schema for the AI scraper asked for name, price, sale price, currency, pack weight in grams, price per 50 g and sale price per 50 g.

The first run looked fine at a glance. Every column was filled. Most values were plausible. You would only see the problem by recomputing each row, which is exactly the check nobody does on a busy day.

What the Model Wrote in the Price per 50 g Column

We went through all 69 yarn rows by hand against the live pages.

What the column held Rows
The price, copied. Right only because the pack is 50 g 18
The price, copied, although the pack is not 50 g 22
About price ÷ 5, a formula nobody asked for 11
The other printed price on the page (the before-price or the sale price) 6
The shop's own printed price per kg 3
Other numbers from the page 3
Empty, because price or pack was missing 6
A real, correct calculation 0

The sale price per 50 g was worse. On 49 rows with no sale at all, the model filled it with a copy of the price. So on a chart, most products looked like they were on sale at their normal price.

We rewrote the schema twice. Once with the formula and an example in the field description, once with an instruction at page level to calculate. Neither changed the pattern. The model copies numbers it can see. It does not produce a number that is not on the page.

The Pack Weight Was Also Wrong, for Different Reasons

The per-50 g value needs a correct pack weight, and that was wrong on 25 of 67 links. These errors were not arithmetic. They were reading errors, and they had causes you can recognise:

  • Another number on the page. A free-shipping threshold of 699 became the pack weight. A pattern note, "about 450 g for a sweater", became 450.
  • Several weights on one page. A product page listing several colourways with different ball sizes gave the weight of a colourway that was not the linked one. That happened on 6 links.
  • Units. Shops that store weight in kilograms gave 0.05 and 0.1, and the model wrote those as grams.
  • Nothing found. On 4 pages the price and the weight were in the page, but not where the model looked, for example inside microdata or written by a script.
  • The wrong product. A collection page with many yarns gave a different yarn.

One of the 25 was not the model's fault at all: the customer's file was out of date, and the shop's 226 g was right.

Why This Happens

An extraction model is trained to find a value in text and return it in a field. That is reading. A price per unit is a derived value: it is not printed on the page, so there is nothing to find, and the model fills the field with the closest thing it can find. The output format hides the difference. A copied number and a calculated number look the same in a JSON field.

You can see the same thing in other LLM work. Models are good at pulling "2.60" and "100 g" out of messy text. They are unreliable at doing the division in the same step, and they give no signal when they skip it.

The Split That Works

We changed the rule for every scraper. The AI copies what is printed, with no maths and no unit changes. Every calculation is an after-scrape rule, which is plain code that does the same thing on every row.

The raw schema now has: name, pack_grams, price, sale_price, currency, availability. Every field description says "copy as printed".

The rules then do the rest:

  1. Take a number from text reads the number in front of a unit word, such as "Wool 100 g" in a product name, and can convert kg to g. We use it where the weight is only in text.
  2. Price per unit computes price ÷ pack_grams × 50. It skips the row when the price or the pack is missing, so an unknown stays empty instead of becoming 0.
  3. Convert currency turns the result into EUR with the ECB rate of the day, and writes the rate and its date next to it.
  4. Replace values from a list maps each shop's stock wording to IN_STOCK, OUT_OF_STOCK, BACKORDER or UNKNOWN.

After the change the arithmetic was right on every row that had a price and a pack, 100% in the first run and 100% in the final setup.

What a Text-Cleaning Rule Cannot Do

Before the "Price per unit" and "Take a number from text" rules existed, we tried to get by with the text-cleaning rule we had, which keeps only the characters you choose. It is useful for codes. It is dangerous for numbers. We tested it on real pack texts:

Input Keep Output Right?
100 g digits 100 Yes
0.05 kg digits, dot, comma 0.05 The digits, but still kg
100m/50g digits 10050 No, metres and grams glued together
3.53 oz digits 353 No, the decimal point is gone
49,95 kr. digits, dot, comma 49,95. No, a trailing dot and a comma decimal

This is why the number rule reads the number in front of a named unit and ignores the other digits on the line.

Use Less AI Where a Structured Source Exists

The second fix was to use the AI less. Many of the shops had a structured source with the weight in grams already as a number: Shopify's variant JSON, WooCommerce's Store API, or itemprop microdata in the HTML. For those links we replaced the AI scraper with an API or microdata scraper. Under 10% of the links that return data still use AI, because no structured source on those shops has the weight.

We choose the source in this order: long-run reliability first, then scale, then cost. Structured sources win on all three. On several live pages the AI step failed and returned no data, while a scraper without the AI step read the same pages at the cheapest fetch tier. And on the first Shopify links we moved, the JSON run cost about a ninth of the AI run. More on finding these sources in how to find the hidden JSON API behind a shop page.

A Checklist for AI Extraction

If you use LLM extraction for prices, this is what we would do from day one:

  • Ask only for values that are printed on the page. No totals, ratios, conversions or "per" values.
  • Put every calculation in a rule you can read, and test the rule on sample values before you save it.
  • Leave a derived value empty when an input is missing. Never 0.
  • Recompute a sample of rows by hand after the first run. A filled column is not a correct column.
  • Check that the sale price is lower than the regular price, and treat "equal" as no sale.
  • If a structured source exists, prefer it over AI for the fields it has.

AI extraction is still the right tool for pages with no structure at all, and it saved us a lot of selector writing on the links where nothing else worked. We just stopped asking it to do sums. If you want to try the same split in your own account, the data scraper takes a URL and suggests the fields, and the MCP walkthrough shows how to have Claude set up the rules for you. For the general trade-offs of AI in scraping, see AI-powered web scraping.

Paste any URL — ScrapeWise handles the anti-bot

Managed infrastructure that adapts when sites change. No proxies, no code, no per-request fees.

Not ready to sign up? See 12 real Amazon rows, 972 columns →

97% accuracy on Amazon benchmarks · no credit card · book a 15-min call →

FAQ

Frequently asked questions

LLM data extraction accuracy: questions answered

An extraction model copies values it can see on the page. A price per unit is not printed, so the model filled the field with the nearest number, usually the price itself. It never did the division.

Only values that are printed on the page, copied as they appear: name, price, sale price, currency, pack size and stock wording. Calculations and unit changes belong in rules.

They are fixed steps that run on every scraped row, such as converting currency, mapping stock wording, cleaning codes, taking a number from text and rescaling a price to a fixed pack size.

It was wrong on 25 of 67 links. It read other numbers such as a free shipping threshold, took the weight of another colourway, or wrote kilograms as grams.

When a page has no API, no JSON-LD and no microdata. In our yarn project that was under 10% of the links, and there the AI copies raw values only.