Blog

Should you block AI crawlers? A decision guide for leaders

"Should we block AI?" usually reaches a leadership meeting as one question. It's really three or four, because "AI crawlers" covers several kinds of software doing very different jobs. Block the wrong one and you vanish from the answers your customers are reading. Here's how to make the call, kind by kind.

Start by splitting the question

The major AI companies each run several agents, and each has its own name so you can treat them differently.

KindWhat it doesExamples
Training crawlersCollect pages to train future AI modelsGPTBot, ClaudeBot, CCBot, Google-Extended (a control name, not a separate crawler)
AI search crawlersIndex pages so AI search can cite and link to youOAI-SearchBot, Claude-SearchBot, PerplexityBot
Assistant fetchersFetch a page in real time because a person asked about itChatGPT-User, Claude-User, Perplexity-User
Browser agentsDrive a real browser for a person: search, fill in forms, add to cartUsually look like an ordinary browser

Our robots.txt guide has the longer list and the exact names. The point for decision-makers is simpler: you can say yes to some and no to others.

The case for each, in business terms

Assistant fetchers: almost always allow

Each visit stands in for a real person asking about you right now: "Is this in stock?", "What's their returns policy?", "Which of these three is cheapest?". Block them and the assistant answers without you, or recommends a competitor. For a store or a software company, this is the closest thing to a customer walking in.

AI search crawlers: allow, unless you also opt out of search

These build the index AI search draws on. They're the AI equivalent of being in Google's index, and they're how you get cited and linked. Few businesses that want search traffic would block Googlebot. The same logic applies here.

Training crawlers: a real choice

This is the one worth debating. Allowing training means your content may shape what AI models know, including about your brand and products. Blocking it keeps your content out of future training sets, and leaves AI search and assistants unaffected, because they use different names.

  • Lean towards blocking if your content is the product: publishers, research, courses, original reviews, anything you license or sell.
  • Lean towards allowing if your content mainly exists to sell something else: product pages, help centers, pricing. Being well understood by AI models is usually worth more to you than the content itself.

Either answer is legitimate. That's why AgentScore reports blocked training crawlers but doesn't take points off for them.

Browser agents: you can't block them by name, so make them work

Browser agents mostly look like a normal browser, so robots.txt rules by name don't reach them. The useful question isn't whether to block them but whether they can complete a purchase or sign-up when they arrive. That's what Ghost Agent tests check.

A quick decision guide

If you are...Assistants and AI searchTraining crawlers
An online storeAllowUsually allow; your product pages are there to be read
A software or services companyAllowUsually allow marketing and docs; your call on anything gated
A publisher or content businessAllow if you want AI citations and referral trafficOften block, or allow only under a licensing agreement
UnsureAllowDecide deliberately; don't inherit someone's copied list

Three mistakes that cost more than the decision

  1. Copying a "block all AI" list. These lists usually include assistant fetchers and AI search crawlers, so you drop out of AI answers along with training.
  2. Deciding in robots.txt and forgetting the CDN. Bot protection, firewall rules and one-click "block AI bots" settings can override a friendly robots.txt. The agent never gets far enough to read it. See bot protection that lets AI agents through.
  3. Trusting the name. Anyone can claim to be ChatGPT or Googlebot. robots.txt is a request, not a lock: well-behaved agents follow it and scrapers ignore it. Stopping abusive traffic, including impostors using trusted names, has to happen at your CDN, by verifying who's really asking. See how to tell if an AI crawler is real.

One more nuance: some operators treat an assistant fetch, made because a person asked about a specific page, like a person clicking a link, and say robots.txt may not apply to it. If you truly need to stop those visits, that's a CDN rule, not a robots.txt line. Most businesses want those visits most of all.

What to do this month

  1. Agree a written policy for each kind of agent: one line each, signed off by marketing, e-commerce and whoever owns your content rights.
  2. Have someone write it into robots.txt. Allowing assistants and search while opting out of training looks like this:
    User-agent: GPTBot
    User-agent: ClaudeBot
    User-agent: CCBot
    User-agent: Google-Extended
    Disallow: /
    
    User-agent: *
    Allow: /
    
    Sitemap: https://northwind.example/sitemap.xml
  3. Ask whoever runs your CDN or bot protection to confirm its settings match the policy, and that it checks agents are genuine rather than trusting their names.
  4. Run AgentScore. It reads your robots.txt as the AI agents we track would, and requests your pages as several AI assistants to see whether your bot protection lets them through. Here's how that works.
  5. Measure what happens next. Count the agents that visit and the visits and sales AI assistants send you, and revisit the policy every quarter.

The default for most businesses is simple: let the agents that bring customers in, make a deliberate choice about training, and enforce both where it actually counts.

← All blog Test your site with AgentScore →

Can AI agents use your site?

Get your free AgentScore in under a minute. No sign-up needed.