Give your AI agent live web data via MCP· Lesson 6 of 6

Guardrails, cost control and untrusted content

Live web access turns an agent into something that can spend money and read text written by strangers. Neither is a reason not to do it. Both are reasons to put limits in before you need them.

  • 11 min read
  • No account needed

Everything up to here has been about making the agent capable. This lesson is about the two things that change the moment it is: it can now spend your money without a human in the loop, and it can now read text that someone else wrote with the intent of being read by a model.

Neither of those is exotic. Both are cheap to handle in advance and genuinely unpleasant to handle after the fact.

Cost: where it actually goes

Three separate meters run at once, and people usually only watch one.

  • Fetches. Every page the agent asks for is a request against your data provider. An agent in a retry loop can make several hundred in a minute without anything looking wrong.
  • Tokens. Tool results go into context. A tool that returns raw HTML can cost more in tokens per call than the fetch did, and it does it on every turn the result stays in the conversation.
  • Repetition. Agents re-check. Without caching, asking "and what about the other three" refetches the first one too. This is the quiet one.

Four limits worth having before you need them

  1. 1

    A hard spend cap, enforced server-side

    Not a dashboard alert — a limit that refuses the request. Alerts tell you about the money after it is gone. Put the cap where the agent cannot route around it, which means in your server or your provider account, not in the prompt.

  2. 2

    A per-conversation call budget

    Most legitimate sessions make a handful of tool calls. If one makes forty, something is looping. Cut it off and return an error the model can read, so it stops rather than retries.

  3. 3

    A short cache on identical requests

    Even sixty seconds removes most of the duplicate-fetch waste, and for scheduled price data a far longer window is usually correct. Return the capture time with the cached result so the staleness is visible.

  4. 4

    Scoped credentials

    A read-only key for a read-only agent. If the only thing an agent is meant to do is read a feed, a key that can also delete scrapers is an unnecessary blast radius.

How to hold that boundary in practice

  • Label it in the payload. Wrap scraped text in a field that marks it as external content, and say in the system prompt that anything inside it is untrusted.
  • Separate reading from acting. The tool that fetches a page should not be the tool that sends an email or writes to your database. Two tools, and a human or an explicit rule between them.
  • Confirm anything with an effect. Reading is cheap to get wrong. Sending, publishing, purchasing and deleting are not. Those get a confirmation step, even when it is annoying.
  • Do not let page content choose the next URL. If the agent follows a link it found in untrusted text, the destination was chosen by whoever wrote that page.
  • Return parsed fields, not prose. A tool that returns {price, currency, in_stock} carries almost no room for an instruction. One that returns the full page description carries plenty.

What gets a confirmation and what does not

ActionReversible?Default
Read a price or a feedYesRun freely
Summarise or compare rowsYesRun freely
Write to an internal sheet or tableMostlyRun, log it
Send a message, email or alert to a personNoConfirm first
Change a listed price or publish anythingNoConfirm first
Delete data or spend moneyNoConfirm first

The thing people get wrong about logging

Log the tool calls, not just the answers. When an agent gives a wrong answer three weeks from now, the question will be whether it fetched the wrong page, got a stale cache, or had good data and reasoned badly — and those three have completely different fixes. Without a record of which tools ran with which arguments, you cannot tell them apart, and you will end up rewriting a prompt to fix a data problem.

A line per call with the tool name, arguments, duration and result size is enough. It costs nothing and it is the difference between debugging and guessing.

Worked example: how a harmless exchange costs real money

Three meters run at once, and the expensive case is the one where they move together. An agent with no per-task cap, asked a broad question against a five-competitor feed of forty rows each, where a row is roughly 300 tokens once it is in context. Four hundred fetches for a question that needed two hundred, and a final turn carrying ten times the context of the first — both for the same reason: nothing told the agent it already had the rows.

TurnFetchesRows in contextTokens in context
1. One competitor4040about 12,000
2. "And the other four"160200about 60,000
3. Re-fetches instead of reusing turn two200400about 120,000
Across the exchange400—about 120,000 in the last turn alone

What usually goes wrong

None of these are reasons not to give an agent live data. They are the things to put in before you need them.

  • Alerting on spend instead of refusing it. An alert tells you about money that has already gone; only a hard cap that returns an error stops the next call.
  • Capping the daily total and nothing else. A single runaway conversation can spend a day's budget in four minutes, which makes a per-task cap the more important of the two.
  • Counting fetches and ignoring tokens. Rows re-read across turns cost nothing in fetches and a great deal in context, as row three above shows.
  • Giving a reading tool and an acting tool the same trust. Anything that writes, sends, buys or deletes belongs behind an explicit confirmation and well away from anything that merely reads.
  • Treating scraped text as part of the conversation. A product description is text written by a stranger; if it reaches the model looking like an instruction you have handed a third party a turn in your prompt. Wrap it, label it as data, and say so in the system prompt.
  • Logging the answers and not the calls. When an output is wrong you will have no way to separate a bad row from a bad summary, and you will spend a week rewriting prompts to fix a data problem.

That is the course. If you have not built the feed yet, course one is the other half of this — picking competitors, matching listings and catching the day a run quietly halves.

Course one: competitor price monitoring

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.