Once a server is connected, the model does not read your code. It reads a list: each tool's name, its description, and a JSON schema of its parameters. That is the whole interface. Everything you know about what the tool does that is not written in those three places is invisible.
This is why two servers that do the same job can behave completely differently. One gets called at the right moment with sensible arguments. The other sits unused, or gets called with the wrong thing and returns an error the model then apologises for. The difference is almost never the implementation.
The four questions a description has to answer
Write the description for a competent colleague who has never seen your system and cannot ask you anything.
- What does this return? Not what it does internally — what comes back. "Returns the current listed price, currency and stock status for one product URL" beats "scrapes a product page".
- When should it be used instead of answering directly? Say it explicitly: "Use this whenever the user asks about a current price. Do not answer from memory; listed prices change daily."
- When should it not be used? Boundaries prevent the expensive failure mode of a tool being called on everything. "This handles one URL at a time. For a whole catalogue, use list_monitored_products."
- What does an argument look like? One real example in the description does more than a paragraph of prose. "url: the full product page URL, e.g. https://www.example.com/p/12345".
The same tool, written two ways
Both of these are valid. Only one gets called when it should.
| Field | Weak version | Version that works |
|---|---|---|
| Name | fetch | get_product_price |
| Description | Fetches data from a URL. | Returns the current listed price, currency, availability and the time it was checked, for a single retailer product page. Use this any time the user asks what something costs right now. Do not answer price questions from memory. |
| Parameter | input (string) | url (string): full product page URL, e.g. https://www.example.com/p/12345 |
| Empty result | Returns [] with no explanation. | Returns a reason string: "page_blocked", "no_price_found" or "not_a_product_page". |
Fewer tools, more clearly separated
There is a strong temptation to expose everything your API can do. Resist it. Every extra tool is another line the model has to read and another chance to pick the wrong one, and the failure mode of too many similar tools is not an error — it is a plausible-looking call to the nearly-right one.
If two tools share most of their description, they should probably be one tool with a parameter. If a tool has eleven optional parameters, it is probably two tools. The test is whether you can state in one sentence when to use each, without using the word "or".
Make results readable, not just correct
The output goes into the model's context as text. Shape it accordingly.
- Return units and currency inside the payload. A bare 24.99 is ambiguous and the model will guess.
- Include when the data was captured. A timestamp in the result is what lets an agent say "as of this morning" instead of implying it is live to the second.
- Keep it small. A 200KB blob of raw HTML crowds out everything else in the context window and the model will usually extract the wrong field from it anyway. Return the five fields you parsed, not the page.
- Say so when there is nothing. An empty array is indistinguishable from a successful check that found no change. A short reason code is not.
- Fail loudly. An error the model can read — "rate limit reached, retry in 60 seconds" — produces sensible behaviour. A silent empty result produces a confident wrong answer.
Worked example: the same result, three payloads
The description decides whether a tool gets called. The payload decides whether the answer is worth anything. All three of these carry the correct price for the same product at the same moment, and only the third lets the model say something a colleague could check.
| What comes back | What the agent can honestly say | |
|---|---|---|
| Raw | a fragment of markup containing 199,00 € | Something about 199, with the decimal comma and the currency position both guessed at |
| Thin | { "price": 199 } | "It is 199." No currency, no date, no source, and nothing to cite |
| Usable | { "price": 199.00, "currency": "EUR", "captured_at": "2026-03-14T06:14:00Z", "availability_text": "In stock", "source_url": "https://…/p/12345" } | "€199.00, in stock, read at 06:14 UTC on 14 March" — with a link, and with the age of the figure visible to the model rather than only to you |
What usually goes wrong
Every item here describes a tool that works perfectly and is used badly, or not at all.
- Writing the description for a human reviewer. The model is the reader. "Returns product data" is a sentence a person understands and a model cannot act on.
- Leaving out the negative case. A description that never says when not to use the tool will see it used for everything, including the questions it answers wrongly.
- Shipping one tool per endpoint. Forty tools means forty near-identical descriptions, and the model is choosing between them on wording alone.
- Returning HTML. Every character of markup is tokens you pay for and a parsing job the model does unreliably.
- Omitting units and currency. 199 is not a price; it is a number that happens to be near one.
- Testing the function rather than the description. The unit test passes, the agent never calls the tool, and no test you have can see that.
Why this lands harder with web data than with most tools
A calendar tool either returns your events or errors. A web data tool has a third state that looks like success: it ran, it returned, and the numbers are wrong because the page layout changed, or you got a localised version of the site, or the listed figure excludes shipping.
Nothing in the protocol protects you from that. The protection is in the payload design — return the source URL, the capture time and a confidence or status field alongside the number, so the model has something to be cautious with. An agent that can see "status: price_from_cache, 3 days old" will hedge. An agent handed a bare number will not.
The principles are easier to see against something concrete. Next, wiring a real price feed into an agent end to end.
Lesson 5: giving an agent a scraper