Ask an assistant what a specific product costs on a specific site today. You will get an answer. It will be formatted confidently, it will be in the right currency, it will be approximately the right magnitude, and there is a good chance it is fabricated.
Understanding exactly why is the difference between fixing it and adding tools at random until something works. There are four causes, and from the outside they look identical.
The four causes
- 1
1. No access at all
The model has no tool that can fetch a page, so it does the only thing it can: produces the most plausible continuation. Training data contained thousands of product pages, so it knows roughly what a price for this kind of product looks like, and it writes one. This is not lying — nothing in the model's situation distinguishes reporting from inferring. The fix is a tool, and this is the only one of the four that adding a tool actually fixes.
- 2
2. It has a tool and did not use it
Far more common than people expect, and much more frustrating. The tool is registered, the agent answers anyway. Usually the description did not make it obvious that this tool was for this question, or the model judged the answer to be something it already knew. You diagnose this by looking at the trace: if there is no tool call, the problem is the description, not the data.
- 3
3. It used the tool and the tool failed quietly
The fetch returned a bot-wall page, a cookie consent interstitial, or a 200 response containing no product at all. If the tool returns that as a string and says nothing about it having failed, the model will do its best with it — which sometimes means extracting a number from the wrong page and sometimes means falling back on its own guess. A tool that cannot fail loudly turns an infrastructure problem into a hallucination.
- 4
4. It used the tool and the data was stale
A cached result from a crawl, a search snippet from a page indexed three weeks ago, or a feed that last ran before the promotion ended. The answer is real data and it is wrong, which is the hardest case to spot because every part of the system reports success.
Diagnosing which one you have
Look at the trace before changing anything. Each cause has a different fix and three of the four fixes do nothing for the others.
| What the trace shows | Cause | Fix |
|---|---|---|
| No tool call at all, direct answer | 1 or 2 | If no tool exists, add one. If one exists, rewrite its description. |
| Tool called, returned quickly, looks empty | 3 | Make the tool return an explicit error instead of an empty success |
| Tool called, plausible content, wrong number | 3 or 4 | Check what page was actually fetched, and when |
| Tool called, correct data, wrong answer | Neither | The result format is confusing the model — see lesson four |
Why "just tell it not to make things up" does not work
Adding "do not guess prices, say you do not know" to a system prompt helps a little and does not solve it. The reason is that from inside the generation process there is no reliable signal distinguishing a fact retrieved from a fact reconstructed. Both arrive the same way. An instruction to only report retrieved facts is an instruction the model cannot fully verify it is following.
What does work is structural: make the data path the only path. If the answer must cite a field from a tool result, and the tool result is absent, there is nothing to cite and the failure becomes visible rather than fluent. Design for a loud absence rather than a confident fallback.
Three properties a live-data tool needs
Everything else in this course is downstream of these.
- It fetches at the moment it is called, and the result carries the timestamp of that fetch so the model and the user can both see how fresh it is
- It fails loudly. A blocked request, an empty page or a missing field returns an explicit error, never an empty success that reads as "no price".
- It returns the specific field, not the page. Handing a model fifty kilobytes of HTML and hoping is a worse version of the problem you started with.
Worked example: reading one trace instead of guessing
An agent answers €189.00 when the page says €229.00. The four causes are indistinguishable from the chat window and take about ninety seconds to separate in the trace. Work down the list and stop at the first "no".
- 1
Was any tool called at all?
If the trace shows no tool invocation, the answer came out of the model's weights and the price is a reconstruction of something that was true at training time. That is cause one, and the fix is configuration rather than prompting.
- 2
Was the right tool called?
A trace showing a web search call and no fetch call is cause two. Search returned a review article from last year quoting €189.00, the model read it, and the citation will look entirely respectable. Nothing failed anywhere.
- 3
Did the call return rows?
Look at the response body, not the status. A 200 carrying an empty rows array is cause three, and it is the dangerous one: the model receives nothing, has nothing to say, and closes the gap. An empty array and a missing field both read as absence of news rather than absence of data.
- 4
How old is the row it did get?
If the response carried a capture time eleven days old, the tool worked perfectly and the data is the problem. That is cause four. The agent has no way of knowing eleven days is too old for a price unless the row says when it was read and the description told it to care.
- 5
Only now change something
The four causes have four different fixes — add a tool, narrow a description, make the empty result loud, stamp and surface the capture time. Editing the system prompt addresses none of them, which is why it is usually the first thing tried and almost never the thing that works.
What usually goes wrong
Four of these five produce an answer that is confident, well-formatted and wrong.
- Adding "do not make things up" to the prompt. The model cannot distinguish retrieving from reconstructing, so the instruction has nothing to act on.
- Handing an agent a general web search tool and calling it a price feed. Search answers what is generally true; a price is what one page says right now.
- Returning an empty array to mean "no data". Return an explicit failure with a reason, because silence is the input most likely to be filled in.
- Leaving the capture time out of the payload. If the row does not carry when it was read, no amount of prompting will make the agent cautious about an eleven-day-old price.
- Letting one tool claim to cover everything. A description that promises every retailer will get the tool called for the one it cannot do, where it fails in exactly the quiet way above.
The honest limit
None of this makes an agent reliable about the live web in general. It makes it reliable about the specific pages you decided to give it access to. An agent with a scraper for twelve named competitors answers questions about those twelve well and is exactly as unreliable as before about everything else.
That is a feature if you treat it as one. Scope the tool narrowly, say in its description what it does and does not cover, and the model will tell the user it cannot answer rather than inventing something. Scope it as "get any price from any website" and you have built a very expensive way to produce the same confident fiction.
Enough theory. Next: actually connecting a server and proving the tools are there.
Lesson 3: connecting an MCP server