A modern extraction API does not have a fixed response shape. You tell it what a product looks like to you, and it goes and finds those things on the page. That is enormously more flexible than a fixed endpoint, and it moves a responsibility onto you that fixed endpoints never did: if you describe the fields badly, you get bad fields, and nothing anywhere will tell you that is what happened.
The schema is a JSON Schema document. Each field has a name, a type, and usually a description. The names are yours — call the price field price or unit_cost or whatever your downstream job already expects, because renaming it later means touching every consumer.
A minimal product schema
Six fields, which is more than most people need and fewer than most people declare.
{
"type": "object",
"properties": {
"product_title": { "type": "string" },
"price": { "type": ["number", "null"] },
"currency": { "type": ["string", "null"] },
"availability": { "type": ["string", "null"] },
"brand": { "type": ["string", "null"] },
"sku": { "type": ["string", "null"] }
},
"required": [
"product_title", "price", "currency",
"availability", "brand", "sku"
]
}Note that every property is also in required, including the ones that are allowed to be null. That is not a contradiction — see below.
The required list is what makes a field appear
This is the single most expensive thing to learn by accident, so here it is plainly: listing a field in required is what makes the extractor emit it. A property that is declared but not required can be omitted entirely, and when it is omitted you do not get a null — you get no key at all, on every row, forever.
The intuition most developers bring is the validation intuition: required means this must not be missing, so leave optional fields out and let them be absent when the page does not have them. That reasoning produces a feed with four columns when you asked for nine, and the symptom looks exactly like "the site does not publish brand" rather than "we never asked for brand".
The fix is to put every field in required and allow null in its type. You are then saying two different things in the two places: the type says this value may legitimately be absent, and the required list says the key must be present regardless. That is the combination that gives you a stable column set, which is what a warehouse table needs.
Type choices that save a conversion later
Small decisions, each of which costs an hour downstream if you get it wrong.
- Numbers should be numbers, with no currency symbol and no thousands separator. A price that arrives as the string "1.299,00 €" is now your parsing problem, and the comma means different things in different markets.
- Use ["number", "null"] rather than ["string", "null"] for anything you will do arithmetic on, even if the page shows it with a unit.
- Keep currency as its own field rather than inferring it from the domain. Multi-market storefronts exist and they will catch you out.
- Availability is better as the raw string the page published than as a boolean you guessed. "Ships in 2–3 weeks" is not in stock and it is not out of stock.
- Always carry the source URL on the row. Every single data question you will be asked begins with "where did this come from".
Declare fewer fields than you think you need
There is a real cost to a wide schema. More fields means more for the extractor to look for, which means more latency and, on some pricing models, more cost per page. It also means more columns that can be empty, and an empty column is not neutral — it is something a colleague will eventually interpret as a zero.
A better pattern is to scrape only the raw values and compute anything derived afterwards. Unit price, price in a common currency, discount percentage: none of those should be fields the extractor hunts for. They are arithmetic on fields you already have, and doing them after collection means a currency rate change does not require a re-run.
There is a mechanical reason for this too. If a derived value is declared as a scraped field, a post-processing rule that wants to write to that same name usually cannot — the scraped column occupies it. So the derived column stays empty and the rule looks broken when actually it was blocked.
Worked example: the same schema, two required arrays
This is the single behaviour that catches most people, and it is worth seeing side by side. Both schemas declare the same five properties. The only difference is what is listed as required — and that difference decides whether a column exists at all. The left-hand column is how a spreadsheet ends up with a header row that moves between runs; the right-hand one is a stable column set with honest empties.
| Field | required is ["name", "price"] | required is all five, types allow null |
|---|---|---|
| name | "Acme Widget 500g" | "Acme Widget 500g" |
| price | 19.99 | 19.99 |
| sale_price, no promotion on the page | key absent from the row entirely | null |
| pack_size, not printed on the page | key absent from the row entirely | null |
| currency, present on the page | sometimes there, sometimes not | "EUR" |
| Columns in the resulting table | varies row by row | five, every row, every run |
What usually goes wrong
Schema mistakes do not raise errors. They produce a table that is a little bit wrong in a way that is hard to see.
- Declaring a field and leaving it out of the required array. It is described, it is understood, and it is not emitted — the quietest failure in this whole course.
- Writing instructions into the field description. A description says what the field is, not what to do with it; arithmetic, conversions and clean-up all belong downstream.
- Declaring a field whose name collides with something computed later. A scraped unit_price column blocks the rule meant to produce one, and the rule then simply never runs.
- Asking for thirty fields on the first pass. Each one is another thing that can be wrong, and twenty-five of them will never be read by anybody.
- Typing a number as a string because the page shows "500 g". Keep the raw text in its own field and let the number field be a number or null, otherwise you are parsing on every read forever.
- Scaling to ten thousand pages without reading one complete row. A missing column is obvious in a single row and invisible in a summary count.
Verify with one page before you run ten thousand
Point the configuration at a single known URL and look at the row. Not the row count — the row. Check that every key you declared is present, that the price is a number, that the currency is what you expected, and that the title is the product rather than the category.
If a column is missing, the schema is the first place to look, not the site. If a column is present but null on a page where you can see the value with your own eyes, that is a genuine extraction problem and worth reporting. Those two failures look identical in a dashboard and have completely different fixes.
Next: why the call that starts a run does not return your data, and how to work with that properly.
Lesson 4: asynchronous runs and polling