# Ghost Agent Labs: every guide and blog post
> Ghost Agent Labs makes websites work for AI agents. It tests whether AI agents can read, navigate, and complete tasks on a website, and shows how to fix what breaks. It is a Ghost Inspector product.
Key facts:
- Who it's for: teams whose website drives revenue, especially e-commerce, growth and SEO, and QA and engineering teams.
- What it does: shows which AI crawlers, assistants and browser agents visit a site, using server and CDN data rather than a JavaScript tag. It also tests journeys such as checkout and sign-up with real AI agents, and scores agent readiness with prioritized fixes.
- What it isn't: a bot blocker. It helps the agents a site wants to succeed, and works alongside existing bot protection.
- Price: AgentScore, the agent-readiness test, is free. The score and top issues need no signup; an email address unlocks the full list of fixes. The Ghost Agent Labs platform is in early access with design partners, and self-serve plans are coming with the public launch.
- Company: built by the team behind Ghost Inspector (https://ghostinspector.com/), the automated browser testing service.
- Our scanner or Ghost Agents visiting your site: email bot@ghostagentlab.com, or see the AgentScore bot and Ghost Agent pages below.
- Markdown: every blog post and article is also available as Markdown by adding `index.html.md` to its URL. llms-full.txt has the full text of all of them in one file.
This file has the full text of all 42 guides and 20 blog posts, guides first. For a short guide to the site, see https://ghostagentlab.com/llms.txt.
---
## What is agent readiness? A complete guide
> How well AI agents can get into your site, understand it, find their way around and complete tasks: what each part means, what good looks like, and where to start.
Published October 8, 2026 by Ghost Agent Labs · Start here · https://ghostagentlab.com/articles/agent-readiness/
### Key takeaways
- Agent readiness is whether AI agents can get into your site, understand it, find their way around and complete tasks for customers.
- An agent that can't use your site doesn't complain or call support; it moves on and recommends a site it can use.
- Start by checking your robots.txt, asking how your bot protection treats AI agents, and fixing the top three issues AgentScore lists.
Agent readiness is how well AI agents can get into your website, understand it, find their way around it, and complete the tasks people send them to do. This guide explains each part, what good looks like, and where to start.
### Why it matters now
AI assistants increasingly visit websites on a person's behalf: to answer a question, compare products, check a policy, or buy something. Each of those visits stands in for a customer. If the agent can't use your site, it doesn't complain or call support. It moves on to a site it can use, and recommends that one instead.
Agent readiness is not the same as SEO, though they overlap. SEO is about being found and ranked. Agent readiness is about what happens next: whether an agent that has found you can actually do something useful with your site.
Not every agent wants the same thing. A crawler collecting pages for AI training is very different from an assistant checking your returns policy for a customer, or a browser agent trying to check out. See [the AI agents that visit your website](https://ghostagentlab.com/articles/kinds-of-ai-agents/), and keep the [agent readiness glossary](https://ghostagentlab.com/articles/agent-readiness-glossary/) handy for the terms in this guide.
### The four parts of agent readiness
[AgentScore](https://ghostagentlab.com/agentscore/) measures agent readiness out of 100, across four categories. The weights reflect what stops agents most often in practice.
| Category | The question | Points |
| --- | --- | --- |
| Access | Can agents get in? | 25 |
| Readability | Can agents understand your pages? | 25 |
| Navigability | Can agents find their way around? | 20 |
| Task completion | Can an agent actually get the job done? | 30 |
#### 1. Access: can agents get in?
Before anything else, the agent has to be allowed through the door. Most access failures are accidental: a robots.txt rule copied from a template, or bot protection that challenges every visitor that isn't a person.
- **robots.txt** allows AI assistants and AI search agents. You can still block training crawlers separately. See [robots.txt for AI agents](https://ghostagentlab.com/articles/robots-txt-ai-agents/).
- **Bot protection** doesn't block AI agents that a person would get through, and can tell real agents from impostors. See [bot protection and CAPTCHAs](https://ghostagentlab.com/articles/bot-protection-ai-agents/) and [how to tell if an AI crawler is real](https://ghostagentlab.com/articles/verify-ai-crawlers/).
- **No CAPTCHA or challenge page** on arrival for ordinary pages.
- **An llms.txt file** points agents to your most important pages. See [how to write an llms.txt file](https://ghostagentlab.com/articles/llms-txt/).
#### 2. Readability: can agents understand your pages?
Many agents read the HTML your server sends and never run JavaScript. If your prices, product details or policies only appear after scripts load, those agents see an empty page.
- **Key content is in the HTML**, through server-side rendering or static generation, not only added by JavaScript. See [why AI agents can't see JavaScript-only content](https://ghostagentlab.com/articles/javascript-content-ai-agents/).
- **Structured data** (schema.org) describes products, prices, availability and your business. See [product data AI shopping agents can read](https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/).
- **Clear titles, headings and image descriptions** on every important page.
- **A valid sitemap** that lists the pages you want found.
#### 3. Navigability: can agents find their way around?
Browser agents find their way around a page much as a screen reader does: by the names and roles of buttons, links and form fields. Much of agent readiness here is simply good accessibility. See [buttons, links and forms agents can use](https://ghostagentlab.com/articles/agent-friendly-buttons-forms/).
- **Every button, link and form field has a name** an agent can read. An icon with no label is invisible.
- **Cookie banners and pop-ups** can be dismissed and don't cover the content.
- **Menus open on click or focus**, not only on hover.
- **Page landmarks** such as header, navigation, main and footer let agents find what they need quickly.
#### 4. Task completion: can an agent get the job done?
This is the part that matters most, and the only way to measure it is to try. A site can pass every technical check and still lose agents at the size selector or the checkout button.
- **Online stores:** an agent can search, find a product, choose options and add it to the cart. See [how to make checkout work for AI shopping agents](https://ghostagentlab.com/articles/agent-ready-checkout/).
- **Software companies:** an agent can find pricing and reach the sign-up form.
- **Service businesses:** an agent can find prices, availability and a way to book.
The free AgentScore scan doesn't test this yet: it marks Task completion “not tested” and works out your score from the other three categories. [Ghost Agents](https://ghostagentlab.com/ghost-agent/) test it today, running the journeys you choose on a schedule, with a step-by-step replay of each run, and alerting you when one breaks.
### How to measure your agent readiness
1. **Run AgentScore.** It's free, takes under a minute, and checks Access, Readability and Navigability. Task completion isn't part of the free scan yet.
2. **Look at your agent traffic.** Your analytics probably can't see agents, because most never run the JavaScript tag. Server or CDN logs can, and Ghost Agent Labs reads them for you. See [how to measure agent traffic](https://ghostagentlab.com/articles/measure-ai-agent-traffic/).
3. **Test your money journeys.** Checkout, sign-up and booking are the journeys where an agent failure costs you a customer.
### A 30-minute starter checklist
> 1. Open `/robots.txt` and make sure it doesn't block AI assistants or AI search agents.
> 2. Ask whoever runs your CDN or bot protection how it treats verified AI agents.
> 3. View the source of a product or pricing page (not the inspector) and check the price is in the HTML.
> 4. Publish an llms.txt file with links to your key pages.
> 5. Tab through your header with the keyboard. If you can't reach the menu or the cart, neither can many agents.
> 6. Run [AgentScore](https://ghostagentlab.com/agentscore/) and fix the top three items it lists.
### Common questions
#### Won't letting agents in mean letting scrapers in?
No. Agent readiness is about letting in the agents you want, and the main ones can be verified. Keep blocking unknown and abusive bots; just make sure your rules don't catch AI assistants that are bringing you customers.
#### Do I have to let AI companies train on my content?
No. Training crawlers and the agents that fetch pages for people use different names, so you can block one and allow the other. Our [robots.txt guide](https://ghostagentlab.com/articles/robots-txt-ai-agents/) shows how.
#### Is this just accessibility?
There's a big overlap, especially in navigation, and work on one helps the other. But agent readiness also covers access rules, machine-readable data, and whether an agent can finish a real task.
### Sources and further reading
- [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html) (IETF)
- [Product](https://schema.org/Product) (schema.org)
- [Sitemaps XML format](https://www.sitemaps.org/protocol.html) (sitemaps.org)
- [The /llms.txt file](https://llmstxt.org/) (llmstxt.org)
---
## The AI agents that visit your website, and what each one wants
> Training crawlers, AI search crawlers, assistant fetchers, browser agents and tool-using agents: what each does, what it needs from your site, and how to treat it.
Published October 9, 2026 by Ghost Agent Labs · Start here · https://ghostagentlab.com/articles/kinds-of-ai-agents/
“AI bots” isn't one kind of visitor. Some collect pages to train models, some build AI search, some fetch a page because a person just asked about you, and some open your site in a browser and click through it. This guide explains the five kinds, what each one needs from your site, and how to treat each.
### Why the difference matters
Each kind of agent does a different job for a different company, and brings you a different kind of value. A training crawler takes your content and may never send anyone back. An assistant fetching your returns policy is answering a customer's question right now. A browser agent on your checkout page is trying to place an order.
Treat them all as “bots” and you get one of two bad outcomes: block everything and disappear from AI answers and purchases, or allow everything without knowing what you've agreed to. The rest of this guide gives you the vocabulary to make a separate decision for each.
### The five kinds at a glance
| Kind | Why it visits | How it reads your site | What it's worth to you |
| --- | --- | --- | --- |
| Training crawler | To collect pages for training AI models | Raw HTML, in bulk | Indirect at best |
| AI search crawler | To index pages so AI search can cite and link to them | Mostly raw HTML | Visibility in AI answers |
| Assistant fetcher | A person asked an assistant about a specific page or question | Usually raw HTML, one page at a time | A customer asking about you right now |
| Browser agent | A person asked an assistant to do something, like book or buy | A real browser: renders the page, clicks and types | A task, often a sale, in progress |
| Tool-using agent | To call tools your site offers directly, through MCP or an API | Structured answers from your tools, not pages | The fastest, most reliable route to a task |
### 1. Training crawlers
Training crawlers collect large numbers of pages to train AI models. Examples include `GPTBot`, `ClaudeBot` and `CCBot`. They don't act for a particular person, and a visit from one doesn't mean anyone is looking for you.
- **What they need:** nothing special. They read your HTML like a search crawler.
- **How to control them:** robots.txt. They use their own names, so you can block them without blocking AI search or assistants. See [robots.txt for AI agents](https://ghostagentlab.com/articles/robots-txt-ai-agents/).
- **The decision:** whether your content may be used for training is a business and legal choice, not a readiness one. AgentScore reports blocked training crawlers but doesn't take points off for them.
### 2. AI search crawlers
AI search crawlers, such as `OAI-SearchBot`, `Claude-SearchBot` and `PerplexityBot`, build the index that AI search and assistants draw on when they answer questions and cite sources. They work much like Googlebot does for traditional search.
- **What they need:** your key content in the HTML your server sends, clear titles and headings, structured data, and a sitemap. Many don't run JavaScript. See [why AI agents can't see JavaScript-only content](https://ghostagentlab.com/articles/javascript-content-ai-agents/).
- **How to control them:** robots.txt, and your CDN or bot protection, which has to let them through.
- **Why allow them:** if they can't read you, AI answers about your category are written from your competitors' pages.
### 3. Assistant fetchers
When someone asks ChatGPT, Claude or Perplexity “what's the return window at this store?” or “compare these two products”, the assistant may fetch the pages it needs there and then. Those requests use names like `ChatGPT-User`, `Claude-User` and `Perplexity-User`.
- **What they need:** fast, readable pages with the facts in plain text: prices, stock, shipping and returns. See [shipping, returns and FAQ pages AI assistants can quote](https://ghostagentlab.com/articles/policy-pages-ai-assistants/).
- **How to control them:** some operators treat these fetches like a person clicking a link and say robots.txt may not apply. Stopping them completely means a rule at your CDN or server.
- **Why allow them:** every one of these visits stands in for a customer asking about you. These are usually the visits you want most.
### 4. Browser agents
Browser agents go further: they open your site in a real browser and use it the way a person would, to search, choose a size, fill in a form or reach checkout. Several AI assistants now offer an agent mode that does this on a person's behalf, and businesses are starting to use them too.
- **What they need:** buttons, links and form fields with clear names, pop-ups they can close, no CAPTCHA in the way, and a checkout that doesn't force an account. They find their way around much as a screen reader does. See [buttons, links and forms agents can use](https://ghostagentlab.com/articles/agent-friendly-buttons-forms/) and [how to make checkout work for AI shopping agents](https://ghostagentlab.com/articles/agent-ready-checkout/).
- **How to recognize them:** often you can't from the name alone, because many use an ordinary browser user agent. Some operators now sign their requests with [Web Bot Auth](https://ghostagentlab.com/articles/verify-ai-crawlers/) so sites can tell who they are.
- **Why they matter:** they're the agents that complete tasks, so they're where agent readiness turns into revenue, and where a broken button or a bot challenge costs you an order.
### 5. Tool-using agents
Instead of reading pages, some agents call tools a site offers them directly: “search products”, “check stock”, “add to cart”. These tools are offered through the Model Context Protocol (MCP), WebMCP in the browser, or agentic commerce protocols for checkout.
- **What they need:** a set of well-described, safe tools, starting with read-only ones such as search and product details.
- **Why it's worth planning for:** tools are faster and more reliable than clicking through pages, and few sites offer them yet. See [MCP and WebMCP](https://ghostagentlab.com/articles/mcp-webmcp-for-websites/) and [agentic commerce protocols explained](https://ghostagentlab.com/articles/agentic-commerce-protocols/).
- **Where to start:** most businesses get more from fixing the first four kinds first. Tools build on readable pages and clean product data, rather than replacing them.
### How to see which ones visit you
Your web analytics probably shows almost none of this. Most agents never run the JavaScript tag analytics depends on, and the ones that do are often filtered out as bots. To see them, look at your server or CDN logs, where every request is recorded with its user agent and IP address. Then check that the visitors claiming to be well-known agents really are: names are easy to fake.
1. **Count by kind, not by bot.** Group visits into the five kinds above, so you can see how many customer-driven fetches you get compared with crawls.
2. **Verify the big names.** Check claimed agents against the IP ranges or signatures their operators publish. See [how to tell if an AI crawler is real](https://ghostagentlab.com/articles/verify-ai-crawlers/).
3. **Watch for errors.** A rise in blocked or failed requests from one kind of agent usually means a rule changed. See [finding the errors AI agents hit](https://ghostagentlab.com/articles/agent-errors-in-logs/).
Ghost Agent Labs does this for you from your CDN or server logs. See [how to measure agent traffic](https://ghostagentlab.com/articles/measure-ai-agent-traffic/).
### A sensible starting policy
> - **Training crawlers:** your choice. Block them by name in robots.txt if you don't want your content used for training.
> - **AI search crawlers and assistant fetchers:** allow them, in robots.txt and at your CDN, and make sure your key content is in the HTML.
> - **Browser agents:** allow verified ones through bot protection, especially on product, cart and checkout pages, and test that they can finish your key journeys.
> - **Tool-using agents:** plan for them once the basics work, starting with read-only tools.
> - **Everyone else:** keep blocking unknown and abusive bots, and anything that claims a trusted name but fails verification.
For a fuller decision guide by business type, see [should you block AI crawlers?](https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/) To see how your site treats each kind today, run [AgentScore](https://ghostagentlab.com/agentscore/): its robots.txt check reads your rules as 21 AI agents would.
---
## Agent readiness glossary: the terms you'll hear, in plain English
> Plain-English definitions of the terms behind agent readiness, from assistant fetchers and Web Bot Auth to structured data, the accessibility tree, MCP and AgentScore.
Published October 9, 2026 by Ghost Agent Labs · Start here · https://ghostagentlab.com/articles/agent-readiness-glossary/
Agent readiness comes with a lot of new words, and some old ones used in new ways. This glossary explains the terms you'll meet in our guides, in AgentScore reports and in conversations with your developers and vendors, in plain English and with what each one means for your business.
### The agents
**AI agent.** Software that uses an AI model to do something on the web for a person or a company: read a page, answer a question, compare products or complete a task. In our guides, “agent” covers everything from crawlers to agents that shop.
**Training crawler.** A bot that collects pages to train AI models, such as `GPTBot` or `ClaudeBot`. Blocking it doesn't stop AI assistants visiting for a person.
**AI search crawler.** A bot that indexes pages so AI search and assistants can cite and link to them, such as `OAI-SearchBot` or `PerplexityBot`.
**Assistant fetcher.** An AI assistant fetching a page in real time because a person asked it something, such as `ChatGPT-User` or `Claude-User`. Each visit stands in for a customer.
**Browser agent.** An agent that opens your site in a real browser and clicks, types and scrolls like a person, to book, sign up or buy. Also called a computer-use agent or autonomous agent.
**User agent.** The name a visitor sends with every request to say what it is, such as a browser version or `GPTBot`. Easy to fake, so it shouldn't be trusted on its own.
For how these differ and how to treat each, see [the AI agents that visit your website](https://ghostagentlab.com/articles/kinds-of-ai-agents/).
### Getting in: access
**robots.txt.** A file at the root of your site that tells crawlers and agents which pages they may visit. Well-behaved agents follow it; scrapers ignore it. See [robots.txt for AI agents](https://ghostagentlab.com/articles/robots-txt-ai-agents/).
**Bot protection.** Services such as Cloudflare, Akamai or DataDome that block or challenge automated visitors. Set too strictly, they turn away AI assistants as well as scrapers. See [bot protection and CAPTCHAs](https://ghostagentlab.com/articles/bot-protection-ai-agents/).
**CDN.** Content delivery network: the service that sits in front of your site and delivers it quickly around the world. It's often where bot rules live, and where agent traffic is recorded. See [setting up your CDN for AI agents](https://ghostagentlab.com/articles/cdn-settings-ai-agents/).
**CAPTCHA or challenge.** A test that tries to tell people from bots, such as picking out images or waiting on a “checking your browser” page. Agents can't pass them, so one on arrival or at checkout ends the visit.
**Rate limit.** A cap on how many requests a visitor can make in a given time. Fair limits stop scrapers without turning away agents. See [rate limits for AI agents](https://ghostagentlab.com/articles/rate-limits-ai-agents/).
**Verification.** Checking that a visitor claiming to be a known agent really is, using the proof its operator publishes: IP ranges, reverse DNS or a signature. See [how to tell if an AI crawler is real](https://ghostagentlab.com/articles/verify-ai-crawlers/).
**Impostor (spoofed agent).** A request that uses a trusted agent's name but fails verification. Usually a scraper trying to get past bot protection.
**Forward-confirmed reverse DNS.** A verification check that an IP address's hostname belongs to the operator, and that the hostname points back to the same address.
**Web Bot Auth.** A new standard in which agents sign their requests cryptographically, so sites can prove who sent them. Our own scanner and Ghost Agents sign theirs. See [why we sign every request](https://ghostagentlab.com/blog/signed-requests/).
**llms.txt.** A short Markdown file at `/llms.txt` that tells AI what your business does and links to your key pages. An emerging standard. See [how to write an llms.txt file](https://ghostagentlab.com/articles/llms-txt/).
**Guest checkout.** Letting shoppers buy without creating an account. An agent buying for someone can't easily sign up for them. See [guest checkout](https://ghostagentlab.com/articles/guest-checkout-ai-agents/).
### Understanding pages: readability
**Raw HTML.** The page exactly as your server sends it, before any JavaScript runs. It's all that many crawlers and assistant fetchers ever see. To look at it, use your browser's “View page source”, not the inspector.
**Server-side rendering (SSR) and static generation.** Ways of building pages so the content is already in the raw HTML, rather than added afterwards by JavaScript. See [why AI agents can't see JavaScript-only content](https://ghostagentlab.com/articles/javascript-content-ai-agents/).
**Structured data (schema.org, JSON-LD).** Machine-readable facts in your page code, such as a product's name, price and stock, written in the shared schema.org vocabulary, usually as JSON-LD. Agents can read it without guessing. See [product data AI shopping agents can read](https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/).
**XML sitemap.** A file listing the pages you want found, so crawlers and agents don't depend on following links. See [XML sitemaps for AI agents](https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/).
**Canonical URL.** The one address you name as the main version of a page that can be reached in several ways. It stops agents treating duplicates as different products. See [canonical URLs](https://ghostagentlab.com/articles/canonical-urls-ai-agents/).
**Alt text.** A short text description of an image. Agents can't see pictures, so this is how they learn what a product photo shows. See [alt text for AI agents](https://ghostagentlab.com/articles/image-alt-text-ai-agents/).
### Finding the way: navigability
**Accessibility tree.** The structured view of a page that screen readers and many browser agents use: its headings, links, buttons and form fields, with their names. If something isn't in it, many agents can't use it. See [what an AI agent sees](https://ghostagentlab.com/blog/what-an-ai-agent-sees/).
**Accessible name.** The name a button, link or field has in the accessibility tree, from its text, label or `aria-label`. An icon-only button with no name is invisible to agents.
**Landmarks.** Labels for the main regions of a page, such as header, navigation, main content and footer, that let agents skip straight to what they need.
**Overlay.** Anything that covers the page, such as a cookie banner, newsletter pop-up or chat window. If an agent can't close it, the journey ends there. See [cookie banners and pop-ups](https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/).
**WCAG.** The Web Content Accessibility Guidelines, the standard for making sites usable by people with disabilities. Much of it helps agents too. See [accessibility work is agent readiness work](https://ghostagentlab.com/blog/accessibility-is-agent-readiness/).
### Getting it done: tasks and commerce
**Journey.** A task a visitor sets out to complete on your site, such as finding a product and reaching checkout, or booking a call. The journeys that make you money are the ones to test first.
**Task completion.** Whether an agent can actually finish a journey. It's the AgentScore category worth the most points, and the only way to measure it is to try. See [how to test your key journeys](https://ghostagentlab.com/articles/test-journeys-with-ai-agents/).
**MCP (Model Context Protocol).** A standard way to give AI agents direct tools for your site, such as search, product details and cart, instead of making them click through pages.
**WebMCP.** A proposal for offering those same tools from inside your web pages, so a browser agent can call them while it's on your site. See [MCP and WebMCP](https://ghostagentlab.com/articles/mcp-webmcp-for-websites/).
**Agentic commerce.** Purchases that an AI agent makes for a shopper. Protocols such as the Agentic Commerce Protocol, the Universal Commerce Protocol and AP2 aim to make that safe for buyer, store and payment provider. See [agentic commerce protocols explained](https://ghostagentlab.com/articles/agentic-commerce-protocols/).
### Measuring it
**Agent readiness.** How well AI agents can get into your site, understand it, find their way around it and complete tasks on it. See [what is agent readiness?](https://ghostagentlab.com/articles/agent-readiness/)
**AgentScore.** Our free score out of 100 for agent readiness, across Access (25 points), Readability (25), Navigability (20) and Task completion (30). The free scan marks Task completion “not tested” and scores the other three. See [how AgentScore works](https://ghostagentlab.com/blog/how-agentscore-works/).
**Check.** One thing AgentScore tests, such as “Buttons and links have names”, reported as a pass, warning or fail, with who usually fixes it.
**Ghost Agent.** A real AI agent we send through a journey on your site, on a schedule, to see whether agents can complete it. You get a step-by-step replay of each run, and an alert when a journey breaks. See [Ghost Agent](https://ghostagentlab.com/ghost-agent/).
**Run and pass rate.** A run is one attempt by one agent at a journey. AI agents don't behave the same way every time, so journeys are tried several times, and the pass rate is the share of runs that succeed.
**Flaky.** A journey where some runs pass and some fail. Agents can do it, but not reliably, which usually points to a slow page, a pop-up that appears sometimes, or an unclear control.
**Agent traffic.** Visits from AI agents of every kind. Most never run your analytics tag, so it's measured from server or CDN logs. See [how to measure agent traffic](https://ghostagentlab.com/articles/measure-ai-agent-traffic/).
**AI referral.** A person who arrives at your site by clicking a link in an AI assistant's answer. Unlike agent visits, these do show up in analytics. See [tracking visits and sales from AI assistants](https://ghostagentlab.com/articles/ai-referral-traffic/).
**Crawls per AI referral.** How many pages an AI company's agents fetched for every visitor its assistant sent you. A low number means you get traffic back for what they read.
To see where your own site stands on each of these, run [AgentScore](https://ghostagentlab.com/agentscore/). It's free and takes under a minute.
---
## robots.txt for AI agents: who to allow and who to block
> The three kinds of AI agent, the robots.txt names each one uses, a sensible default, and the most-specific-group rule that catches most people out.
Published October 8, 2026 by Ghost Agent Labs · Access · https://ghostagentlab.com/articles/robots-txt-ai-agents/
### Key takeaways
- You can block AI training crawlers and still let AI search and assistants, which answer customers' questions about you, read your site.
- A crawler follows only the most specific group that names it, so adding a group for one agent can quietly drop your other rules.
- Open your robots.txt today and check it doesn't block AI search or assistant agents, including a leftover rule that blocks everything.
Your robots.txt file may be blocking the AI assistants that send you customers, while doing nothing to stop the bots you actually wanted to keep out. This guide explains the three kinds of AI agent, which robots.txt names each one uses, and how to write rules that let the right ones in.
### Three kinds of AI agent
"AI bots" isn't one thing. The major AI companies each run several agents with different jobs and different names, and the right policy for each is different.
| Kind | What it does | Examples (robots.txt name) |
| --- | --- | --- |
| Training crawlers | Collect pages to train AI models | `GPTBot`, `ClaudeBot`, `CCBot`, `Bytespider`, `meta-externalagent` |
| AI search crawlers | Index pages so AI search can cite and link to them | `OAI-SearchBot`, `Claude-SearchBot`, `PerplexityBot`, `Amazonbot` |
| Assistant fetchers | Fetch a page in real time because a person asked about it | `ChatGPT-User`, `Claude-User`, `Perplexity-User`, `meta-externalfetcher`, `DuckAssistBot`, `MistralAI-User` |
Some names aren't crawlers at all. `Google-Extended` and `Applebot-Extended` are control tokens: Google and Apple crawl with Googlebot and Applebot as usual, and these names let you say whether that content may be used for their AI models. Blocking them doesn't affect normal search.
Lists like this change as companies launch new agents, so check each operator's documentation for the current names.
### A sensible default
For most businesses, the agents that matter most are AI search and assistant fetchers, because they answer people's questions about you and link back to you. Training is a separate decision about your content, and it's yours to make.
This robots.txt allows everything, which is also what you get with no robots.txt at all:
```
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml
```
If you'd rather not have your content used to train AI models, but still want to appear in AI search and assistant answers, block only the training names:
```
# Allow AI search and assistants; opt out of AI training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /
User-agent: *
Disallow: /cart
Disallow: /account
Allow: /
Sitemap: https://www.example.com/sitemap.xml
```
Listing several `User-agent` lines above one set of rules applies those rules to all of them.
### The rule that catches most people out
A crawler follows only the most specific group that matches its name, and ignores the rest. If there's a group for `GPTBot`, GPTBot ignores everything under `User-agent: *`.
So in the example above, the training crawlers don't inherit the `/cart` and `/account` rules; they don't need them, because they're blocked from everything. But the reverse mistake is common:
```
User-agent: *
Disallow: /checkout
Disallow: /admin
User-agent: OAI-SearchBot
Allow: /
```
Someone added that second group to "make sure" OpenAI's search crawler gets in. Instead, they've told it that `/checkout` and `/admin` are open, because it no longer reads the `*` group. If you add a group for a specific agent, copy into it every rule you want it to follow.
### Other common mistakes
- **A leftover `Disallow: /`.** Staging sites are often launched with a robots.txt that blocks everything. Check yours today.
- **Blocking every AI name from a copied list.** Long "block all AI" lists usually include the assistant fetchers and AI search crawlers too, which removes you from AI answers as well as training.
- **Blocking the pages agents need.** Disallowing `/products` or `/pricing` for all bots, or for `*`, to save crawl budget also hides them from agents.
- **Relying on robots.txt to block bad bots.** robots.txt is a request, not a lock. Well-behaved crawlers follow it; scrapers ignore it. Abusive traffic has to be stopped at your CDN or firewall.
- **Forgetting your CDN.** A firewall rule or bot setting that challenges "AI crawlers" overrides a friendly robots.txt. The agent never gets far enough to read it.
### A note on assistant fetchers
When a person asks an assistant to read a specific page, some operators treat that fetch like a person clicking a link rather than a crawl, and say robots.txt rules may not apply to it. If you need to stop those visits completely, that has to happen at your CDN or server, not in robots.txt. For most businesses, though, these are the visits you want most: each one is a customer asking about you.
### How to check your robots.txt
1. Open `https://yourdomain.com/robots.txt` and read it top to bottom. Note every group and which names it applies to.
2. For each AI search crawler and assistant fetcher in the table above, work out which group it follows, remembering the most-specific-group rule, and whether that group blocks your important pages.
3. Run [AgentScore](https://ghostagentlab.com/agentscore/). Its robots.txt check reads your rules as 21 AI agents would and tells you which assistants and AI search agents are blocked from your home page. Blocked training crawlers are reported but don't cost you points, because that's a legitimate choice.
robots.txt is only the first door. Next, make sure your bot protection lets the real agents through, and can tell them apart from impostors using their names. See [how to tell if an AI crawler is real](https://ghostagentlab.com/articles/verify-ai-crawlers/).
### Sources and further reading
- [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html) (IETF)
- [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) (OpenAI)
- [Does Anthropic crawl data from the web, and how can site owners block the crawler?](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) (Anthropic)
- [List of Google's common crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers) (Google)
---
## How to tell if an AI crawler is real
> Anyone can claim to be GPTBot or Googlebot. How to verify AI agents with published IP ranges, reverse DNS and signed requests, and what to do when there's no proof.
Published October 8, 2026 by Ghost Agent Labs · Access · https://ghostagentlab.com/articles/verify-ai-crawlers/
### Key takeaways
- Anyone can send a request claiming to be GPTBot or Googlebot, so trusting the name alone lets impostors straight in.
- Treat requests you can't verify as unknown rather than fake, because blocking all of them can shut out legitimate agents.
- Ask whether your CDN already verifies bots, and make sure the AI assistants and AI search agents you want are allowed.
Anyone can send a request that says it's GPTBot or Googlebot. To let real AI agents in while keeping impostors out, you need to check where a request actually came from. There are three ways to do it, and the right one depends on the operator.
### Why user agents can't be trusted
Every request names its sender in the `User-Agent` header, and that name is a promise, not a proof. Scrapers routinely borrow the names of well-known crawlers, because many sites let those crawlers through without a challenge. If your bot rules trust the name alone, you're letting in anyone who copies it. If they distrust it, you're blocking real agents along with the fakes.
Verification resolves that. It answers one question: did this request really come from the company it names?
### Method 1: published IP ranges
Several operators publish the IP addresses their crawlers use, as machine-readable JSON files. If a request claims to be one of their agents and comes from an address in the list, it's real.
| Operator | Agents covered |
| --- | --- |
| Google | Googlebot, special-case crawlers and user-triggered fetchers (separate lists) |
| Microsoft | Bingbot |
| Apple | Applebot |
| OpenAI | GPTBot, OAI-SearchBot and ChatGPT-User (one list each) |
| Perplexity | PerplexityBot and Perplexity-User (one list each) |
Each operator links its list from its crawler documentation. A few practical points:
- **Refresh the lists often.** Ranges change. Fetch them at least daily, and don't treat a request as fake because it's missing from a list that's days old.
- **Match the list to the agent.** An OpenAI address in the GPTBot list doesn't verify a request claiming to be ChatGPT-User.
- **Handle IPv6.** Several lists include IPv6 ranges.
### Method 2: forward-confirmed reverse DNS
Some operators instead promise that their crawlers' IP addresses resolve to a host name on their own domain. The check has two steps, because reverse DNS on its own can be faked by whoever controls the IP address:
1. **Reverse lookup:** look up the host name for the request's IP address, and check it ends in the operator's domain.
2. **Forward lookup:** look up that host name's IP addresses, and check the original IP is among them.
1. A request arrives
User agent says Googlebot, from IP address `66.249.66.1`
2. 1. Reverse lookup
`66.249.66.1` → `crawl-66-249-66-1.googlebot.com`
Ends in googlebot.com
3. 2. Forward lookup
`crawl-66-249-66-1.googlebot.com` → `66.249.66.1`
Points back to the same address
4. Verified
The request is from Google. If either step fails, it isn’t, whatever its user agent says.
*Forward-confirmed reverse DNS, using Google’s own example. The second step matters because whoever controls an IP address can make its reverse lookup say anything.*
Google's own example, from a terminal:
```
$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1
```
The name ends in `googlebot.com` and points back to the same address, so the request is from Google. Domains that operators document for this include:
| Agent | Host name ends in |
| --- | --- |
| Googlebot and other Google crawlers | `googlebot.com`, `google.com` or `googleusercontent.com` |
| Bingbot | `search.msn.com` |
| Applebot | `applebot.apple.com` |
| Amazonbot | `crawl.amazonbot.amazon` |
DNS lookups are slow compared with a page request, so cache the result for each IP address rather than checking on every hit.
### Method 3: signed requests (Web Bot Auth)
The newest method doesn't depend on IP addresses at all. With [Web Bot Auth](https://blog.cloudflare.com/web-bot-auth/), the agent signs each request with a private key and publishes the matching public key on its own domain. Your CDN or server checks the signature, and a valid one proves the sender, whatever network it came from.
That makes it the best fit for browser agents running in the cloud, whose addresses change constantly. It's still early, so check whether your CDN supports it. We explain how it works, and why we sign all of our own agents' requests, in [Why we sign every request our agents send](https://ghostagentlab.com/blog/signed-requests/).
### When there's no proof at all
Not every operator publishes IP ranges, a reverse DNS domain, or signing keys for every agent. For those, you can't prove a request is real or fake; you can only weigh the evidence, such as how it behaves and how fast it requests pages.
Treat "can't tell" as its own answer, not as "fake". Blocking everything you can't verify will block legitimate agents from operators that simply haven't published proof yet.
### Putting it together
1. **Identify:** match the user agent to a known agent and its operator.
2. **Verify:** use the strongest proof that operator offers: a signature, then published IP ranges, then reverse DNS.
3. **Decide:** let verified agents you want through; challenge or block requests that claim a name but fail verification; and handle "can't tell" with ordinary rate limits rather than a hard block.
> **Your CDN may already do this.** Many CDNs and bot-management products maintain lists of verified bots. Check that the AI assistants and AI search agents you want are in the allowed categories, and that "AI crawler" blocking rules aren't catching them by name.
>
> **Ghost Agent Labs does it for you.** It identifies every agent in your server or edge traffic, checks it against published IP ranges and reverse DNS, and flags impostors, so you can see which agents are real before you decide what to allow. Missing or stale data is reported as "can't tell", never as "spoofed". [Start free](https://app.ghostagentlab.com/signup).
### Sources and further reading
- [Verify requests from Google crawlers and fetchers](https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests) (Google)
- [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) (OpenAI)
- [About Applebot](https://support.apple.com/en-us/119829) (Apple)
- [HTTP Message Signatures for automated traffic Architecture (Web Bot Auth draft)](https://datatracker.ietf.org/doc/draft-meunier-web-bot-auth-architecture/) (IETF)
---
## Bot protection and CAPTCHAs: stop blocking the agents you want
> Why bot protection catches AI assistants, how to tell if it's happening, and how to tune your CDN, WAF and CAPTCHAs so good agents get through and bad bots don't.
Published October 8, 2026 by Ghost Agent Labs · Access · https://ghostagentlab.com/articles/bot-protection-ai-agents/
### Key takeaways
- Bot protection often stops the AI assistants answering customers' questions about your products, because those agents look like bots for legitimate reasons.
- Never let traffic through on its name alone, and check that any one-click block AI setting doesn't also block assistants and AI search.
- Keep product, pricing and policy pages open, and save strict challenges and CAPTCHAs for login, sign-up and checkout.
Bot protection exists to keep out scrapers, credential stuffers and fraud. But the same rules often stop the AI assistants trying to answer a customer's question about your products. This guide shows how to tell whether that's happening, and how to tune your defenses so the agents you want get through and the bots you don't still get stopped.
### Why good agents get caught
Bot protection, whether it's a CDN feature, a web application firewall (WAF) or a dedicated product, looks for signs that a visitor isn't a person. AI agents show many of those signs, for entirely legitimate reasons:
- **They say they're bots.** Honest agents announce themselves in their user agent, and a rule that blocks "bots" or "AI crawlers" by name catches them first.
- **They don't run JavaScript.** Many challenges work by running a script in the visitor's browser. Assistant fetchers and AI crawlers usually read only the HTML, so they never pass.
- **They come from data centers.** Agents run in the cloud, and cloud IP addresses score as higher risk than home broadband.
- **Browser agents look automated.** Even agents that drive a real browser move and type differently from people, and fingerprinting tools notice.
- **One-click "block AI" settings.** Several providers offer a single switch to block AI bots. Depending on how it's set up, it can block assistants and AI search along with training crawlers.
### Signs it's happening to you
- AI assistants say they "can't access" your site, or describe it from out-of-date information.
- Your CDN or firewall logs show 403 (forbidden), 429 (too many requests) or challenge responses for agents like ChatGPT-User, Claude-User or Perplexity-User.
- You appear in AI search results much less often than competitors with weaker SEO.
The quickest test is [AgentScore](https://ghostagentlab.com/agentscore/). It requests your home page as a normal browser and as five AI agents, including ChatGPT-User, Claude-User, Perplexity-User and OAI-SearchBot, and compares what comes back. It fails the check if an agent is blocked or challenged while the browser gets through, and warns if an agent gets less than half the content the browser does. Separate checks look for CAPTCHAs on your home page, and test whether agents can open your product, pricing and cart pages.
### How to fix it
#### 1. Allow verified agents, not names
The safe way to let agents in is by verified identity. Most bot management products keep a list of verified bots, checked against the operator's published IP ranges, reverse DNS or signatures, and let you allow categories of them. Allow the categories for AI assistants and AI search, and keep blocking requests that only claim those names. Our guide to [telling real AI crawlers from fakes](https://ghostagentlab.com/articles/verify-ai-crawlers/) explains how verification works.
> **Never allow by user agent alone.** A rule like "if the user agent contains GPTBot, skip all checks" is the first thing scrapers exploit.
#### 2. Separate training from assistants
If you've turned on a "block AI bots" setting, check exactly which agents it covers. If your goal is to opt out of AI training, block only training crawlers, and leave assistant fetchers and AI search crawlers allowed. See [robots.txt for AI agents](https://ghostagentlab.com/articles/robots-txt-ai-agents/) for which names are which.
#### 3. Protect actions, not pages
Abuse mostly targets actions: logging in, creating accounts, applying discount codes, checking gift card balances, submitting payment. It rarely needs protecting against on pages that only display information. Concentrate strict rules and challenges on:
- Login, sign-up and password reset forms
- Checkout submission and payment
- Gift card, coupon and stock-check endpoints that attackers hammer
- Search and APIs, with rate limits rather than outright blocks
Keep product, category, pricing, policy and help pages as open as you safely can. Those are the pages agents need to answer questions about you.
#### 4. Use rate limits instead of blocks
A real assistant fetcher makes a handful of requests because a person asked about you. A scraper makes thousands. A sensible rate limit per IP address or verified agent stops the scraper without punishing the assistant, where a blanket block stops both.
#### 5. Replace puzzle CAPTCHAs on key pages
AI agents can't solve CAPTCHAs, and aren't meant to. If you need a challenge, use invisible or risk-based challenges that only escalate when traffic looks abusive, and keep them off the pages agents need to read. Never put a CAPTCHA in front of ordinary content on arrival.
#### 6. Don't serve agents a different page
Some setups don't block agents outright, but quietly serve them a stripped-down page, a placeholder, or an "enable JavaScript" message. To the agent this is as bad as a block, and it's harder to notice. Agents should get the same content a browser does.
### A quick checklist
> 1. Run [AgentScore](https://ghostagentlab.com/agentscore/) and look at the bot protection, CAPTCHA and key pages checks.
> 2. In your CDN or bot management settings, find the verified bots or AI categories, and confirm AI assistants and AI search are allowed.
> 3. Check that any "block AI" setting covers only the agents you mean to block.
> 4. Search your firewall rules for user agent matches and make sure none allow access by name alone.
> 5. Move strict challenges to login, sign-up and checkout submission, and use rate limits elsewhere.
> 6. Re-run AgentScore to confirm the fix.
To keep an eye on this over time, watch how often agents are blocked in your traffic. See [how to measure AI agent traffic](https://ghostagentlab.com/articles/measure-ai-agent-traffic/).
### Sources and further reading
- [RFC 9110: HTTP Semantics](https://www.rfc-editor.org/rfc/rfc9110.html) (IETF)
- [RFC 6585: Additional HTTP Status Codes](https://www.rfc-editor.org/rfc/rfc6585.html) (IETF)
---
## How to write an llms.txt file
> What llms.txt is, what to put in it, a complete example for an online store, and the mistakes that make it useless to AI agents.
Published October 8, 2026 by Ghost Agent Labs · Access · https://ghostagentlab.com/articles/llms-txt/
### Key takeaways
- An llms.txt file is a short plain-text guide that tells AI agents what your business does and where your key pages are.
- It's a proposal that not every AI company reads, so treat it as one part of agent readiness, not a replacement for clear pages.
- List twenty to forty well-described links to the pages customers ask about, publish it at your domain root, and keep it current.
An llms.txt file is a short Markdown guide to your website, written for AI. It tells an agent what your business does and where the pages that matter are, so it doesn't have to work it out from menus, pop-ups and marketing copy. It takes about half an hour to write, and AgentScore checks for it.
### What llms.txt is
[llms.txt](https://llmstxt.org/) is a proposed standard, first published by Jeremy Howard of Answer.AI in 2024. The idea is simple: put a plain-text Markdown file at `/llms.txt` on your domain, the same way `/robots.txt` lives at the root. Where robots.txt tells crawlers what they may visit, llms.txt tells AI models and agents what is worth reading.
It's useful because AI agents have limited attention. A typical page is mostly navigation, scripts, tracking and layout. An agent asked "what's the return policy at this store?" has to dig through all of that. A good llms.txt hands it the answer's location in one request.
> **An honest caveat:** llms.txt is a proposal, not something every AI company has committed to read. Some coding assistants and documentation tools use it today, and it's cheap to add, but treat it as one part of agent readiness alongside clear pages, structured data and a sitemap, not a replacement for them.
### The format
The file is ordinary Markdown with a light structure, in this order:
1. **An H1 with your name.** This is the only required part.
2. **A blockquote summary**: one or two sentences on what you do and who for.
3. **Optional paragraphs** with anything an agent should know up front, such as where you ship or what you don't sell.
4. **H2 sections of links**, each a Markdown list item: `- [Page name](URL): what it's for`.
5. **An "Optional" section** for links an agent can skip if it's short on space.
### A complete example
Here's an llms.txt for a fictional online store:
```
# Northwind Coffee
> Northwind Coffee roasts and sells specialty coffee beans, ground coffee and
> brewing equipment online, shipping to the US and Canada.
Orders over $40 ship free in the US. We don't sell gift cards or wholesale online;
wholesale enquiries go through the contact page.
## Shop
- [All coffee](https://northwind.example/collections/coffee): Every bean we sell, with roast level, origin and price
- [Subscriptions](https://northwind.example/subscriptions): Recurring deliveries every 2, 4 or 6 weeks; pause or cancel any time
- [Brewing equipment](https://northwind.example/collections/equipment): Grinders, kettles and pour-over kits
## Help and policies
- [Shipping](https://northwind.example/pages/shipping): Rates, delivery times and countries we ship to
- [Returns](https://northwind.example/pages/returns): 30-day returns on unopened items and equipment
- [FAQ](https://northwind.example/pages/faq): Grind sizes, freshness, subscriptions and account questions
- [Contact](https://northwind.example/pages/contact): Email and live chat hours
## Optional
- [Our story](https://northwind.example/pages/about): Who we are and how we source our beans
- [Brewing guides](https://northwind.example/blog/guides): Step-by-step guides for each brewing method
```
You can also see [our own llms.txt](https://ghostagentlab.com/llms.txt).
### What to include
Think about the questions people ask agents about businesses like yours, and link to the page that answers each one.
| Type of site | Link to |
| --- | --- |
| Online store | Main categories, bestsellers, shipping, returns, sizing, FAQ, contact |
| Software company | Product overview, pricing, sign-up, docs, API reference, security, status page |
| Local or service business | Services and prices, booking, locations and hours, service area, contact |
| Publisher | Sections, subscription options, licensing and permissions, editorial policy |
If you offer agents a direct way in, such as an API or an MCP server, list it with a link to its documentation. AgentScore looks for an MCP server mentioned in llms.txt as one of the signals that a site supports agents directly.
### Common mistakes
- **Serving HTML.** Some site builders return your home page or a 404 page for any unknown URL. Open `/llms.txt` in a browser and make sure you see plain text, not a web page.
- **Writing marketing copy.** "The world's most loved coffee experience" tells an agent nothing. Say what you sell, where, and on what terms.
- **Linking to everything.** It's a guide, not a sitemap. Twenty to forty well-described links beats five hundred bare ones. Your sitemap.xml already lists every page.
- **Leaving out descriptions.** The text after each link is what lets an agent pick the right page without opening them all.
- **Letting it go stale.** Dead links and old prices are worse than no file. Review it whenever you change your navigation or policies.
- **Linking to pages agents can't read.** If a linked page only shows content after JavaScript runs, or sits behind a CAPTCHA, the link won't help. AgentScore checks for both.
### How to publish it
1. Write the file in any text editor and save it as `llms.txt`.
2. Upload it so it's served at the root of your domain: `https://yourdomain.com/llms.txt`. On most hosts this means putting it in the same folder as robots.txt. On Shopify, WordPress and other platforms, apps and plugins can serve it for you.
3. Check it's served as plain text with a 200 status, and that your bot protection doesn't challenge AI agents that request it.
4. Run [AgentScore](https://ghostagentlab.com/agentscore/). The "llms.txt guide for AI" check will pass once the file is live.
### What about llms-full.txt?
Some sites, mostly documentation sites, also publish `/llms-full.txt`: the full text of their key pages in one Markdown file, so an agent can load everything at once. It's a community convention rather than part of the proposal. It's worth it for developer docs; for most stores and marketing sites a good llms.txt is enough.
### Sources and further reading
- [The /llms.txt file](https://llmstxt.org/) (llmstxt.org)
---
## MCP and WebMCP: giving AI agents a direct way into your site
> What MCP servers and WebMCP tools are, what they mean for your business, how AgentScore detects them, and how to start safely.
Published October 9, 2026 by Ghost Agent Labs · Access · https://ghostagentlab.com/articles/mcp-webmcp-for-websites/
### Key takeaways
- MCP and WebMCP let AI agents call named actions, such as searching products, instead of clicking through your pages.
- Prices, stock and policies your tools return must match your website exactly, or agents will quote figures your checkout won't honor.
- Ask your commerce platform about MCP support first, and start with read-only tools for search, product details, pricing and policies.
Most AI agents use your website the way a person does: they read pages, click links and fill in forms. That works, but it's slow and it breaks easily. MCP and WebMCP give agents a second way in, a set of named actions such as "search products" or "check order status" that they can call directly. This guide explains both, what they mean for your business, and how to start small.
### Two ways an agent can use your site
Today, an AI agent asked to find a waterproof jacket in a medium on your store has to work your pages. It loads the home page, finds the search box, reads the results, opens a product, works out the size picker, and reads the price. Every step depends on your page layout. A redesigned menu or a new pop-up can stop it.
The alternative is to tell the agent what it can do and let it ask directly. Instead of driving the search box, it calls a `search_products` action with "waterproof jacket, medium" and gets back a clean list of products, prices and stock. That's the idea behind both MCP and WebMCP.
- **Faster.** One request instead of a dozen page loads and clicks.
- **More reliable.** Actions don't change when you redesign the page.
- **More accurate.** You decide exactly what data comes back, so agents quote the right price and stock.
It doesn't replace a usable website. Many agents will keep using your pages for a long time, and the people they hand back to always will. Think of MCP and WebMCP as an extra, faster lane for the agents that support them.
### What MCP is
The [Model Context Protocol](https://modelcontextprotocol.io/) (MCP) is an open protocol for connecting AI applications to outside tools and data. Anthropic introduced it in late 2024, and it has since been adopted by many AI assistants, developer tools and software companies.
An MCP server is a small service that describes a set of **tools** (actions an agent can call, each with a name, a plain-language description and the inputs it expects) and can also offer **resources** (data it can read, such as a catalog or a policy). An AI application connects to the server, reads the list, and calls the tools it needs to complete a person's request.
For a website, a remote MCP server runs alongside your site, usually on your own domain, and talks to the same systems your site does. Typical tools for an online store:
| Tool | What it does | Risk |
| --- | --- | --- |
| `search_products` | Finds products by keyword, category, size or price | Read-only |
| `get_product` | Returns price, variants, stock and delivery estimate for one product | Read-only |
| `get_policy` | Returns shipping, returns or warranty terms | Read-only |
| `get_order_status` | Looks up an order for a signed-in customer | Needs sign-in |
| `add_to_cart` | Builds a cart and returns a link for the person to check out | Changes state |
Software companies might offer tools to look up plans and pricing or search documentation. Service businesses might offer tools to check availability and request a booking.
Commerce platforms are starting to offer MCP servers for the stores they host, so check with yours before building one. If you have a developer team and an existing API, a read-only server is often a small project, because it wraps calls you already make.
### What WebMCP is
WebMCP is an early proposal for bringing the same idea into the web page itself. Instead of running a separate server, your page registers tools with the browser through a new JavaScript API, `navigator.modelContext`. An agent working in that browser can then call the page's tools directly instead of clicking through it.
It's being developed in the open at the W3C's Web Machine Learning Community Group, with engineers from Google and Microsoft among those working on it. Community group work is incubation, not a finished standard, and browsers don't ship WebMCP by default yet. Expect the details to change.
Two features make it interesting for websites:
- **It reuses your page code.** The tool can call the same functions your search box or add-to-cart button already calls.
- **It works with the person's session.** Because it runs in the page, it can use the cart and sign-in the person already has, without separate API credentials.
A tool registered in a page looks roughly like this, guarded so it does nothing in browsers without the API:
```
if ("modelContext" in navigator) {
navigator.modelContext.registerTool({
name: "search_products",
description: "Search Northwind Outdoor products by keyword, size and maximum price.",
inputSchema: {
type: "object",
properties: {
query: { type: "string" },
size: { type: "string" },
maxPrice: { type: "number" }
},
required: ["query"]
},
async execute({ query, size, maxPrice }) {
const results = await searchCatalog({ query, size, maxPrice }); // your existing search
return { content: [{ type: "text", text: JSON.stringify(results) }] };
}
});
}
```
The proposal also describes a declarative form: adding a `toolname` attribute (with a description) to an existing `