# robots.txt for AI agents: who to allow and who to block

> The three kinds of AI agent, the robots.txt names each one uses, a sensible default, and the most-specific-group rule that catches most people out.

Published October 8, 2026 by Ghost Agent Labs · Access · https://ghostagentlab.com/articles/robots-txt-ai-agents/

## Key takeaways

- You can block AI training crawlers and still let AI search and assistants, which answer customers' questions about you, read your site.
- A crawler follows only the most specific group that names it, so adding a group for one agent can quietly drop your other rules.
- Open your robots.txt today and check it doesn't block AI search or assistant agents, including a leftover rule that blocks everything.

Your robots.txt file may be blocking the AI assistants that send you customers, while doing nothing to stop the bots you actually wanted to keep out. This guide explains the three kinds of AI agent, which robots.txt names each one uses, and how to write rules that let the right ones in.

## Three kinds of AI agent

"AI bots" isn't one thing. The major AI companies each run several agents with different jobs and different names, and the right policy for each is different.

| Kind | What it does | Examples (robots.txt name) |
| --- | --- | --- |
| Training crawlers | Collect pages to train AI models | `GPTBot`, `ClaudeBot`, `CCBot`, `Bytespider`, `meta-externalagent` |
| AI search crawlers | Index pages so AI search can cite and link to them | `OAI-SearchBot`, `Claude-SearchBot`, `PerplexityBot`, `Amazonbot` |
| Assistant fetchers | Fetch a page in real time because a person asked about it | `ChatGPT-User`, `Claude-User`, `Perplexity-User`, `meta-externalfetcher`, `DuckAssistBot`, `MistralAI-User` |

Some names aren't crawlers at all. `Google-Extended` and `Applebot-Extended` are control tokens: Google and Apple crawl with Googlebot and Applebot as usual, and these names let you say whether that content may be used for their AI models. Blocking them doesn't affect normal search.

Lists like this change as companies launch new agents, so check each operator's documentation for the current names.

## A sensible default

For most businesses, the agents that matter most are AI search and assistant fetchers, because they answer people's questions about you and link back to you. Training is a separate decision about your content, and it's yours to make.

This robots.txt allows everything, which is also what you get with no robots.txt at all:

```
User-agent: *
Allow: /

Sitemap: https://www.example.com/sitemap.xml
```

If you'd rather not have your content used to train AI models, but still want to appear in AI search and assistant answers, block only the training names:

```
# Allow AI search and assistants; opt out of AI training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /

User-agent: *
Disallow: /cart
Disallow: /account
Allow: /

Sitemap: https://www.example.com/sitemap.xml
```

Listing several `User-agent` lines above one set of rules applies those rules to all of them.

## The rule that catches most people out

A crawler follows only the most specific group that matches its name, and ignores the rest. If there's a group for `GPTBot`, GPTBot ignores everything under `User-agent: *`.

So in the example above, the training crawlers don't inherit the `/cart` and `/account` rules; they don't need them, because they're blocked from everything. But the reverse mistake is common:

```
User-agent: *
Disallow: /checkout
Disallow: /admin

User-agent: OAI-SearchBot
Allow: /
```

Someone added that second group to "make sure" OpenAI's search crawler gets in. Instead, they've told it that `/checkout` and `/admin` are open, because it no longer reads the `*` group. If you add a group for a specific agent, copy into it every rule you want it to follow.

## Other common mistakes

- **A leftover `Disallow: /`.** Staging sites are often launched with a robots.txt that blocks everything. Check yours today.
- **Blocking every AI name from a copied list.** Long "block all AI" lists usually include the assistant fetchers and AI search crawlers too, which removes you from AI answers as well as training.
- **Blocking the pages agents need.** Disallowing `/products` or `/pricing` for all bots, or for `*`, to save crawl budget also hides them from agents.
- **Relying on robots.txt to block bad bots.** robots.txt is a request, not a lock. Well-behaved crawlers follow it; scrapers ignore it. Abusive traffic has to be stopped at your CDN or firewall.
- **Forgetting your CDN.** A firewall rule or bot setting that challenges "AI crawlers" overrides a friendly robots.txt. The agent never gets far enough to read it.

## A note on assistant fetchers

When a person asks an assistant to read a specific page, some operators treat that fetch like a person clicking a link rather than a crawl, and say robots.txt rules may not apply to it. If you need to stop those visits completely, that has to happen at your CDN or server, not in robots.txt. For most businesses, though, these are the visits you want most: each one is a customer asking about you.

## How to check your robots.txt

1. Open `https://yourdomain.com/robots.txt` and read it top to bottom. Note every group and which names it applies to.
2. For each AI search crawler and assistant fetcher in the table above, work out which group it follows, remembering the most-specific-group rule, and whether that group blocks your important pages.
3. Run [AgentScore](https://ghostagentlab.com/agentscore/). Its robots.txt check reads your rules as 21 AI agents would and tells you which assistants and AI search agents are blocked from your home page. Blocked training crawlers are reported but don't cost you points, because that's a legitimate choice.

robots.txt is only the first door. Next, make sure your bot protection lets the real agents through, and can tell them apart from impostors using their names. See [how to tell if an AI crawler is real](https://ghostagentlab.com/articles/verify-ai-crawlers/).

## Sources and further reading

- [RFC 9309: Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html) (IETF)
- [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) (OpenAI)
- [Does Anthropic crawl data from the web, and how can site owners block the crawler?](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) (Anthropic)
- [List of Google's common crawlers](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers) (Google)
