Articles

robots.txt for AI agents: who to allow and who to block

Usually fixed by: SEO or content team · Typical effort: hours

Your robots.txt file may be blocking the AI assistants that send you customers, while doing nothing to stop the bots you actually wanted to keep out. This guide explains the three kinds of AI agent, which robots.txt names each one uses, and how to write rules that let the right ones in.

Three kinds of AI agent

"AI bots" isn't one thing. The major AI companies each run several agents with different jobs and different names, and the right policy for each is different.

KindWhat it doesExamples (robots.txt name)
Training crawlersCollect pages to train AI modelsGPTBot, ClaudeBot, CCBot, Bytespider, meta-externalagent
AI search crawlersIndex pages so AI search can cite and link to themOAI-SearchBot, Claude-SearchBot, PerplexityBot, Amazonbot
Assistant fetchersFetch a page in real time because a person asked about itChatGPT-User, Claude-User, Perplexity-User, meta-externalfetcher, DuckAssistBot, MistralAI-User

Some names aren't crawlers at all. Google-Extended and Applebot-Extended are control tokens: Google and Apple crawl with Googlebot and Applebot as usual, and these names let you say whether that content may be used for their AI models. Blocking them doesn't affect normal search.

Lists like this change as companies launch new agents, so check each operator's documentation for the current names.

A sensible default

For most businesses, the agents that matter most are AI search and assistant fetchers, because they answer people's questions about you and link back to you. Training is a separate decision about your content, and it's yours to make.

This robots.txt allows everything, which is also what you get with no robots.txt at all:

User-agent: *
Allow: /

Sitemap: https://www.example.com/sitemap.xml

If you'd rather not have your content used to train AI models, but still want to appear in AI search and assistant answers, block only the training names:

# Allow AI search and assistants; opt out of AI training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: meta-externalagent
Disallow: /

User-agent: *
Disallow: /cart
Disallow: /account
Allow: /

Sitemap: https://www.example.com/sitemap.xml

Listing several User-agent lines above one set of rules applies those rules to all of them.

The rule that catches most people out

A crawler follows only the most specific group that matches its name, and ignores the rest. If there's a group for GPTBot, GPTBot ignores everything under User-agent: *.

So in the example above, the training crawlers don't inherit the /cart and /account rules; they don't need them, because they're blocked from everything. But the reverse mistake is common:

User-agent: *
Disallow: /checkout
Disallow: /admin

User-agent: OAI-SearchBot
Allow: /

Someone added that second group to "make sure" OpenAI's search crawler gets in. Instead, they've told it that /checkout and /admin are open, because it no longer reads the * group. If you add a group for a specific agent, copy into it every rule you want it to follow.

Other common mistakes

  • A leftover Disallow: /. Staging sites are often launched with a robots.txt that blocks everything. Check yours today.
  • Blocking every AI name from a copied list. Long "block all AI" lists usually include the assistant fetchers and AI search crawlers too, which removes you from AI answers as well as training.
  • Blocking the pages agents need. Disallowing /products or /pricing for all bots, or for *, to save crawl budget also hides them from agents.
  • Relying on robots.txt to block bad bots. robots.txt is a request, not a lock. Well-behaved crawlers follow it; scrapers ignore it. Abusive traffic has to be stopped at your CDN or firewall.
  • Forgetting your CDN. A firewall rule or bot setting that challenges "AI crawlers" overrides a friendly robots.txt. The agent never gets far enough to read it.

A note on assistant fetchers

When a person asks an assistant to read a specific page, some operators treat that fetch like a person clicking a link rather than a crawl, and say robots.txt rules may not apply to it. If you need to stop those visits completely, that has to happen at your CDN or server, not in robots.txt. For most businesses, though, these are the visits you want most: each one is a customer asking about you.

How to check your robots.txt

  1. Open https://yourdomain.com/robots.txt and read it top to bottom. Note every group and which names it applies to.
  2. For each AI search crawler and assistant fetcher in the table above, work out which group it follows, remembering the most-specific-group rule, and whether that group blocks your important pages.
  3. Run AgentScore. Its robots.txt check reads your rules as 21 AI agents would and tells you which assistants and AI search agents are blocked from your home page. Blocked training crawlers are reported but don't cost you points, because that's a legitimate choice.

robots.txt is only the first door. Next, make sure your bot protection lets the real agents through, and can tell them apart from impostors using their names. See how to tell if an AI crawler is real.

← All articles Test your site with AgentScore →

Can AI agents use your site?

Get your free AgentScore in under a minute. No sign-up needed.