[{"data":1,"prerenderedAt":219},["ShallowReactive",2],{"learn-lesson-ai-agent-web-data-mcp-why-agents-get-live-data-wrong":3},{"course":4,"lesson":67,"index":185,"outline":186,"prev":217,"next":218},{"slug":5,"order":6,"level":7,"time":8,"card_text":9,"seo":10,"hero":16,"outcomes":26,"who":35,"syllabus":46,"faq":49,"lessonCount":66},"ai-agent-web-data-mcp",2,"Comfortable editing a config file","6 lessons, about 60 minutes","Your agent is confidently wrong about prices because it has never seen one. What MCP is, how to connect a server, how to design tools a model can actually use, and the guardrails you need before you let it loose.",{"title":11,"description":12,"keywords":13,"og_title":14,"og_description":15},"MCP for AI Agents: A Free 6-Lesson Course on Live Web Data","What MCP is, how to connect a server to Claude, how to design tools an agent can use, and how to give an agent live web data without it inventing prices. Free, ungated.","mcp tutorial, model context protocol, ai agent web data, mcp server claude, give agent live data, mcp tools design, agent web scraping","A free course on giving AI agents live web data via MCP","Six written lessons: the protocol, the failure modes, connecting a server, designing usable tools, and the guardrails.",{"badge":17,"title":18,"subtitle":19,"cta_primary":20,"cta_secondary":23},"Course two","Give your AI agent live web data via MCP","Ask an assistant what a product costs today and you will usually get a number. It is often wrong, and it is always wrong in the same way: the model is reconstructing a plausible price from training data rather than looking at a page. This course is about closing that gap properly — what the Model Context Protocol actually is, how to wire a server into a client, how to design tools a model can use without hand-holding, and what to put in place before an agent spends your money.",{"label":21,"url":22},"Start with lesson one","/learn/ai-agent-web-data-mcp/what-is-mcp",{"label":24,"url":25},"See our MCP skill","/skills",{"title":27,"items":28},"What you will be able to do",[29,30,31,32,33,34],"Explain MCP to a colleague in two sentences without using the word \"ecosystem\"","Tell the difference between a model that does not know something and a model that has been given a tool badly","Connect an MCP server to a client and verify that the tools are actually registered","Write a tool description a model picks correctly on the first attempt","Give an agent the ability to fetch a real price from a real page","Put a spend cap, a rate limit and an injection boundary in place before any of this touches production",{"title":36,"for_title":37,"for":38,"not_title":42,"not_for":43},"Who this is for","Written for",[39,40,41],"Developers building on Claude, ChatGPT or an agent framework who need the agent to see the live web","Technical founders evaluating whether MCP is worth adopting","Data teams who already have an API and are deciding whether to expose it to agents","Not written for",[44,45],"Anyone looking for a no-code agent builder — this assumes a config file and a terminal","Readers who want the full specification; this is the working subset, and the spec is linked where it matters",{"title":47,"intro":48},"The six lessons","Lessons one and two are concepts and cost nothing to read. From three onwards you will want a client installed.",{"badge":50,"title":51,"description":52,"items":53},"FAQ","Before you start","The questions that come up in the first ten minutes.",[54,57,60,63],{"title":55,"description":56},"Do I need to know what MCP is already?","No. Lesson one assumes nothing beyond having used an AI assistant. If you already know what a tool call is, skim it and start at lesson two.",{"title":58,"description":59},"Is this Claude-specific?","MCP is an open protocol and the concepts transfer to any client that implements it. The concrete configuration examples use Claude because that is the client most readers will have in front of them, and the differences elsewhere are mostly about where the config file lives.",{"title":61,"description":62},"Do I need a Scrapewise account?","Only for lesson five, which walks through pointing an agent at a real scraper. Everything else works against any MCP server, including ones you write yourself in an afternoon.",{"title":64,"description":65},"Can I just use a web search tool instead?","Sometimes, and lesson two is explicit about when. Search gives an agent a summary of a page; a scraper gives it the specific field from a specific page. For \"what is the general sentiment on X\" search is better. For \"what does this exact URL charge today\" it is not.",6,{"slug":68,"nav_title":69,"title":70,"summary":71,"time":72,"needs_account":73,"seo":74,"blocks":78,"takeaways":176,"next_step":181},"why-agents-get-live-data-wrong","Why agents get it wrong","Why your agent's answer about a price is wrong","Four distinct failure modes that all look identical from the outside, and how to tell which one you have before you try to fix it.","10 min",false,{"title":75,"description":76,"keywords":77},"Why AI Agents Give Wrong Answers About Live Data","Four reasons an agent's answer about a live price is wrong — no access, bad tool choice, a blocked fetch, and stale cache — and how to diagnose which one you have.","ai agent wrong answers, llm hallucinate prices, agent live data, why llm outdated information, agent tool selection failure",[79,84,100,105,130,135,143,162,171],{"type":80,"paragraphs":81},"prose",[82,83],"Ask an assistant what a specific product costs on a specific site today. You will get an answer. It will be formatted confidently, it will be in the right currency, it will be approximately the right magnitude, and there is a good chance it is fabricated.","Understanding exactly why is the difference between fixing it and adding tools at random until something works. There are four causes, and from the outside they look identical.",{"type":85,"title":86,"items":87},"steps","The four causes",[88,91,94,97],{"title":89,"text":90},"1. No access at all","The model has no tool that can fetch a page, so it does the only thing it can: produces the most plausible continuation. Training data contained thousands of product pages, so it knows roughly what a price for this kind of product looks like, and it writes one. This is not lying — nothing in the model's situation distinguishes reporting from inferring. The fix is a tool, and this is the only one of the four that adding a tool actually fixes.",{"title":92,"text":93},"2. It has a tool and did not use it","Far more common than people expect, and much more frustrating. The tool is registered, the agent answers anyway. Usually the description did not make it obvious that this tool was for this question, or the model judged the answer to be something it already knew. You diagnose this by looking at the trace: if there is no tool call, the problem is the description, not the data.",{"title":95,"text":96},"3. It used the tool and the tool failed quietly","The fetch returned a bot-wall page, a cookie consent interstitial, or a 200 response containing no product at all. If the tool returns that as a string and says nothing about it having failed, the model will do its best with it — which sometimes means extracting a number from the wrong page and sometimes means falling back on its own guess. A tool that cannot fail loudly turns an infrastructure problem into a hallucination.",{"title":98,"text":99},"4. It used the tool and the data was stale","A cached result from a crawl, a search snippet from a page indexed three weeks ago, or a feed that last ran before the promotion ended. The answer is real data and it is wrong, which is the hardest case to spot because every part of the system reports success.",{"type":101,"variant":102,"title":103,"text":104},"callout","warning","Search results are not current prices","The most common accidental version of cause four is giving the agent a web search tool and assuming it has solved live data. A search result is a snippet from an index, built whenever the crawler last visited. For prices, promotions and stock it is routinely days or weeks out of date, and it presents as authoritative text with a real URL attached. Search answers \"what is generally true about this\". It does not answer \"what does this page say right now\".",{"type":106,"title":107,"intro":108,"headers":109,"rows":113},"table","Diagnosing which one you have","Look at the trace before changing anything. Each cause has a different fix and three of the four fixes do nothing for the others.",[110,111,112],"What the trace shows","Cause","Fix",[114,118,122,126],[115,116,117],"No tool call at all, direct answer","1 or 2","If no tool exists, add one. If one exists, rewrite its description.",[119,120,121],"Tool called, returned quickly, looks empty","3","Make the tool return an explicit error instead of an empty success",[123,124,125],"Tool called, plausible content, wrong number","3 or 4","Check what page was actually fetched, and when",[127,128,129],"Tool called, correct data, wrong answer","Neither","The result format is confusing the model — see lesson four",{"type":80,"title":131,"paragraphs":132},"Why \"just tell it not to make things up\" does not work",[133,134],"Adding \"do not guess prices, say you do not know\" to a system prompt helps a little and does not solve it. The reason is that from inside the generation process there is no reliable signal distinguishing a fact retrieved from a fact reconstructed. Both arrive the same way. An instruction to only report retrieved facts is an instruction the model cannot fully verify it is following.","What does work is structural: make the data path the only path. If the answer must cite a field from a tool result, and the tool result is absent, there is nothing to cite and the failure becomes visible rather than fluent. Design for a loud absence rather than a confident fallback.",{"type":136,"title":137,"intro":138,"items":139},"list","Three properties a live-data tool needs","Everything else in this course is downstream of these.",[140,141,142],"It fetches at the moment it is called, and the result carries the timestamp of that fetch so the model and the user can both see how fresh it is","It fails loudly. A blocked request, an empty page or a missing field returns an explicit error, never an empty success that reads as \"no price\".","It returns the specific field, not the page. Handing a model fifty kilobytes of HTML and hoping is a worse version of the problem you started with.",{"type":85,"title":144,"intro":145,"items":146},"Worked example: reading one trace instead of guessing","An agent answers €189.00 when the page says €229.00. The four causes are indistinguishable from the chat window and take about ninety seconds to separate in the trace. Work down the list and stop at the first \"no\".",[147,150,153,156,159],{"title":148,"text":149},"Was any tool called at all?","If the trace shows no tool invocation, the answer came out of the model's weights and the price is a reconstruction of something that was true at training time. That is cause one, and the fix is configuration rather than prompting.",{"title":151,"text":152},"Was the right tool called?","A trace showing a web search call and no fetch call is cause two. Search returned a review article from last year quoting €189.00, the model read it, and the citation will look entirely respectable. Nothing failed anywhere.",{"title":154,"text":155},"Did the call return rows?","Look at the response body, not the status. A 200 carrying an empty rows array is cause three, and it is the dangerous one: the model receives nothing, has nothing to say, and closes the gap. An empty array and a missing field both read as absence of news rather than absence of data.",{"title":157,"text":158},"How old is the row it did get?","If the response carried a capture time eleven days old, the tool worked perfectly and the data is the problem. That is cause four. The agent has no way of knowing eleven days is too old for a price unless the row says when it was read and the description told it to care.",{"title":160,"text":161},"Only now change something","The four causes have four different fixes — add a tool, narrow a description, make the empty result loud, stamp and surface the capture time. Editing the system prompt addresses none of them, which is why it is usually the first thing tried and almost never the thing that works.",{"type":136,"title":163,"intro":164,"items":165},"What usually goes wrong","Four of these five produce an answer that is confident, well-formatted and wrong.",[166,167,168,169,170],"Adding \"do not make things up\" to the prompt. The model cannot distinguish retrieving from reconstructing, so the instruction has nothing to act on.","Handing an agent a general web search tool and calling it a price feed. Search answers what is generally true; a price is what one page says right now.","Returning an empty array to mean \"no data\". Return an explicit failure with a reason, because silence is the input most likely to be filled in.","Leaving the capture time out of the payload. If the row does not carry when it was read, no amount of prompting will make the agent cautious about an eleven-day-old price.","Letting one tool claim to cover everything. A description that promises every retailer will get the tool called for the one it cannot do, where it fails in exactly the quiet way above.",{"type":80,"title":172,"paragraphs":173},"The honest limit",[174,175],"None of this makes an agent reliable about the live web in general. It makes it reliable about the specific pages you decided to give it access to. An agent with a scraper for twelve named competitors answers questions about those twelve well and is exactly as unreliable as before about everything else.","That is a feature if you treat it as one. Scope the tool narrowly, say in its description what it does and does not cover, and the model will tell the user it cannot answer rather than inventing something. Scope it as \"get any price from any website\" and you have built a very expensive way to produce the same confident fiction.",[177,178,179,180],"Four causes — no tool, unused tool, silent tool failure, stale data — look identical from outside. Read the trace first.","A web search tool is not a live price tool. It answers what is generally true, not what a page says now.","Prompting a model not to guess does not work, because it cannot tell retrieval from reconstruction. Make absence loud instead.","Narrow the tool's scope deliberately. An agent that knows what it does not cover will say so.",{"text":182,"label":183,"url":184},"Enough theory. Next: actually connecting a server and proving the tools are there.","Lesson 3: connecting an MCP server","/learn/ai-agent-web-data-mcp/connect-an-mcp-server",1,[187,193,194,199,205,212],{"slug":188,"navTitle":189,"title":190,"summary":191,"time":192,"needsAccount":73},"what-is-mcp","What MCP is","What MCP actually is, in plain terms","The Model Context Protocol described without jargon: what problem it solves, its three primitives, and when it is the wrong tool.","9 min",{"slug":68,"navTitle":69,"title":70,"summary":71,"time":72,"needsAccount":73},{"slug":195,"navTitle":196,"title":197,"summary":198,"time":72,"needsAccount":73},"connect-an-mcp-server","Connecting a server","Connecting an MCP server and proving it works","The config for local and remote servers, the four things that go wrong, and how to verify the tools registered rather than assuming.",{"slug":200,"navTitle":201,"title":202,"summary":203,"time":204,"needsAccount":73},"design-tools-an-agent-can-use","Designing usable tools","Designing tools an agent can actually use","A connected server is not a useful server. The model only sees your tool names, descriptions and parameter schemas, so those three things are the entire user interface. Here is what makes a tool get called correctly and what makes it get ignored.","11 min",{"slug":206,"navTitle":207,"title":208,"summary":209,"time":210,"needsAccount":211},"give-an-agent-a-scraper","Giving an agent a scraper","Giving an agent a real price feed","A worked example. Connect the ScrapeWise MCP server to a client, let the agent read a live scraper's output, and watch where the hand-off between \"the data is right\" and \"the answer is right\" actually breaks.","12 min",true,{"slug":213,"navTitle":214,"title":215,"summary":216,"time":204,"needsAccount":73},"guardrails-cost-and-untrusted-content","Guardrails and cost","Guardrails, cost control and untrusted content","Live web access turns an agent into something that can spend money and read text written by strangers. Neither is a reason not to do it. Both are reasons to put limits in before you need them.",{"slug":188,"navTitle":189,"title":190,"summary":191,"time":192,"needsAccount":73},{"slug":195,"navTitle":196,"title":197,"summary":198,"time":72,"needsAccount":73},1791047866923]