The previous four lessons were about the protocol and the writing. This one is the whole loop with a real feed behind it: a scraper that already runs on a schedule, exposed to an agent, with the verification step that tells you whether to trust what comes back.
The order matters. Build the feed first and check it by hand, then connect it. Connecting an unverified feed to an agent means you now have two things that might be wrong and no way to tell which.
Wiring it up
- 1
Have a scraper that already produces rows
Course one covers this end to end. The short version: a scraper with a product URL list, a schema with the fields you care about, and at least one completed run whose output you have eyeballed. If you have not looked at the rows yourself, do that before going further.
- 2
Create an API key
In the portal, generate a key scoped to what the agent needs. Keep it out of your repository and out of anything you paste into a chat window — put it in the client config file or an environment variable.
- 3
Add the server to your client config
Same shape as lesson three: an entry under mcpServers with the command and the key passed as an environment variable. Restart the client fully afterwards — a reload is not enough in most clients.
- 4
Confirm the tools appear
Open the tool list in your client before asking a question. If the tools are not listed, nothing you ask will reach them, and the model will cheerfully answer from memory instead.
- 5
Ask a question you already know the answer to
Pick a product whose current price you have open in another tab. Ask the agent. Compare. This single step catches market mismatches, stale caches and currency confusion in about ten seconds.
A config entry with the key in the environment
The exact server command is in the docs; the structure is what matters here.
{
"mcpServers": {
"scrapewise": {
"command": "npx",
"args": ["-y", "@scrapewise/mcp"],
"env": {
"SCRAPEWISE_API_KEY": "sw_live_..."
}
}
}
}Restart the client after editing. Then check the tool list before asking anything.
What the agent can now do that it could not before
With a feed behind it, the useful questions stop being "what does this cost" and start being comparative and historical, because the agent can read many rows at once: which of our SKUs are priced above every competitor we track, which competitor moved most this week, which products went out of stock at two retailers on the same day.
Those are questions a human would answer with a pivot table and twenty minutes. The agent answers them in a sentence, and — this is the part that matters — it can show you the rows it used.
Four failures you will hit, and what each looks like
- The agent answers without calling anything. The tool list is empty, or the description does not connect to the question. Check the list first, then the description.
- The numbers are right but old. The scraper has not run since its last schedule. Always surface the capture time in the answer so this is visible rather than silent.
- The agent quotes a price for the wrong market. The source URL resolved to a different country's storefront. Check which URL is on the link list, not which one you typed into your browser.
- The agent summarises confidently over half the rows. A run partially failed and returned fewer products than usual. This is the dangerous one, because the answer looks fine. Course one's lesson on silent failure is the fix.
Worked example: three questions to ask before you trust it
Once the tools register, ask these three in this order, before the agent goes anywhere near a real task. Each one is chosen because its failure is informative rather than merely annoying.
| Ask | A good answer | What a bad answer is telling you |
|---|---|---|
| "List the tools you have and say what each one returns." | The tool names, each with the shape of its result | If it invents a tool, or describes one you never connected, you are in a stale session — restart before going further |
| "What is the price of X?" — a product you checked in a browser one minute ago | Your figure, its currency, and the time it was captured | A near miss, right magnitude and wrong number, almost always means an old row rather than a broken tool |
| "How many rows did that run return, and when did it run?" | A count and a timestamp, both taken from the payload | Vagueness here means the counts are not in the payload at all, and every summary you get after this point is unfalsifiable |
What usually goes wrong
The feed being right and the answer being right are two separate claims, and this is where people stop distinguishing them.
- Connecting a feed you have not verified by hand. When the answer comes back wrong you then have two suspects and no way to separate them.
- Asking an open question first. "How are we priced against the market?" produces a fluent paragraph nobody can check. Start with something that has exactly one right answer.
- Accepting a summary that cites nothing. An agent summarising a half-failed run sounds exactly like an agent summarising a complete one, and the shortfall is invisible unless row counts are in the payload and required in the answer.
- Letting it aggregate silently. "On average you are 4% above market" across 380 of 1,200 SKUs is a different sentence from the same claim across 1,180, and the denominator will not be volunteered.
- Treating the first good answer as proof. One correct response establishes that the path is wired, not that it is reliable. The failure you actually care about arrives on the morning a run half-completes.
Where this stops being a toy
The step up is giving the agent a standing job rather than answering questions. A morning summary of every competitor move above a threshold. A check before a promotion goes out. A Slack message when a tracked product goes out of stock at the retailer you compete hardest with.
All of those are the same tools, called on a trigger instead of by a person. Which is exactly the point at which cost and guardrails stop being theoretical — the subject of the last lesson.
An agent that can fetch is an agent that can spend, and one that reads pages it does not control. Both need limits.
Lesson 6: guardrails, cost and untrusted content