The AI agents that visit your website, and what each one wants
“AI bots” isn't one kind of visitor. Some collect pages to train models, some build AI search, some fetch a page because a person just asked about you, and some open your site in a browser and click through it. This guide explains the five kinds, what each one needs from your site, and how to treat each.
Why the difference matters
Each kind of agent does a different job for a different company, and brings you a different kind of value. A training crawler takes your content and may never send anyone back. An assistant fetching your returns policy is answering a customer's question right now. A browser agent on your checkout page is trying to place an order.
Treat them all as “bots” and you get one of two bad outcomes: block everything and disappear from AI answers and purchases, or allow everything without knowing what you've agreed to. The rest of this guide gives you the vocabulary to make a separate decision for each.
The five kinds at a glance
| Kind | Why it visits | How it reads your site | What it's worth to you |
|---|---|---|---|
| Training crawler | To collect pages for training AI models | Raw HTML, in bulk | Indirect at best |
| AI search crawler | To index pages so AI search can cite and link to them | Mostly raw HTML | Visibility in AI answers |
| Assistant fetcher | A person asked an assistant about a specific page or question | Usually raw HTML, one page at a time | A customer asking about you right now |
| Browser agent | A person asked an assistant to do something, like book or buy | A real browser: renders the page, clicks and types | A task, often a sale, in progress |
| Tool-using agent | To call tools your site offers directly, through MCP or an API | Structured answers from your tools, not pages | The fastest, most reliable route to a task |
1. Training crawlers
Training crawlers collect large numbers of pages to train AI models. Examples include GPTBot, ClaudeBot and CCBot. They don't act for a particular person, and a visit from one doesn't mean anyone is looking for you.
- What they need: nothing special. They read your HTML like a search crawler.
- How to control them: robots.txt. They use their own names, so you can block them without blocking AI search or assistants. See robots.txt for AI agents.
- The decision: whether your content may be used for training is a business and legal choice, not a readiness one. AgentScore reports blocked training crawlers but doesn't take points off for them.
2. AI search crawlers
AI search crawlers, such as OAI-SearchBot, Claude-SearchBot and PerplexityBot, build the index that AI search and assistants draw on when they answer questions and cite sources. They work much like Googlebot does for traditional search.
- What they need: your key content in the HTML your server sends, clear titles and headings, structured data, and a sitemap. Many don't run JavaScript. See why AI agents can't see JavaScript-only content.
- How to control them: robots.txt, and your CDN or bot protection, which has to let them through.
- Why allow them: if they can't read you, AI answers about your category are written from your competitors' pages.
3. Assistant fetchers
When someone asks ChatGPT, Claude or Perplexity “what's the return window at this store?” or “compare these two products”, the assistant may fetch the pages it needs there and then. Those requests use names like ChatGPT-User, Claude-User and Perplexity-User.
- What they need: fast, readable pages with the facts in plain text: prices, stock, shipping and returns. See shipping, returns and FAQ pages AI assistants can quote.
- How to control them: some operators treat these fetches like a person clicking a link and say robots.txt may not apply. Stopping them completely means a rule at your CDN or server.
- Why allow them: every one of these visits stands in for a customer asking about you. These are usually the visits you want most.
4. Browser agents
Browser agents go further: they open your site in a real browser and use it the way a person would, to search, choose a size, fill in a form or reach checkout. Several AI assistants now offer an agent mode that does this on a person's behalf, and businesses are starting to use them too.
- What they need: buttons, links and form fields with clear names, pop-ups they can close, no CAPTCHA in the way, and a checkout that doesn't force an account. They find their way around much as a screen reader does. See buttons, links and forms agents can use and how to make checkout work for AI shopping agents.
- How to recognize them: often you can't from the name alone, because many use an ordinary browser user agent. Some operators now sign their requests with Web Bot Auth so sites can tell who they are.
- Why they matter: they're the agents that complete tasks, so they're where agent readiness turns into revenue, and where a broken button or a bot challenge costs you an order.
5. Tool-using agents
Instead of reading pages, some agents call tools a site offers them directly: “search products”, “check stock”, “add to cart”. These tools are offered through the Model Context Protocol (MCP), WebMCP in the browser, or agentic commerce protocols for checkout.
- What they need: a set of well-described, safe tools, starting with read-only ones such as search and product details.
- Why it's worth planning for: tools are faster and more reliable than clicking through pages, and few sites offer them yet. See MCP and WebMCP and agentic commerce protocols explained.
- Where to start: most businesses get more from fixing the first four kinds first. Tools build on readable pages and clean product data, rather than replacing them.
How to see which ones visit you
Your web analytics probably shows almost none of this. Most agents never run the JavaScript tag analytics depends on, and the ones that do are often filtered out as bots. To see them, look at your server or CDN logs, where every request is recorded with its user agent and IP address. Then check that the visitors claiming to be well-known agents really are: names are easy to fake.
- Count by kind, not by bot. Group visits into the five kinds above, so you can see how many customer-driven fetches you get compared with crawls.
- Verify the big names. Check claimed agents against the IP ranges or signatures their operators publish. See how to tell if an AI crawler is real.
- Watch for errors. A rise in blocked or failed requests from one kind of agent usually means a rule changed. See finding the errors AI agents hit.
Ghost Agent Labs does this for you from your CDN or server logs. See how to measure agent traffic.
A sensible starting policy
- Training crawlers: your choice. Block them by name in robots.txt if you don't want your content used for training.
- AI search crawlers and assistant fetchers: allow them, in robots.txt and at your CDN, and make sure your key content is in the HTML.
- Browser agents: allow verified ones through bot protection, especially on product, cart and checkout pages, and test that they can finish your key journeys.
- Tool-using agents: plan for them once the basics work, starting with read-only tools.
- Everyone else: keep blocking unknown and abusive bots, and anything that claims a trusted name but fails verification.
For a fuller decision guide by business type, see should you block AI crawlers? To see how your site treats each kind today, run AgentScore: its robots.txt check reads your rules as 21 AI agents would.