Articles

How to tell if an AI crawler is real

Usually fixed by: Bot protection or CDN admin

Anyone can send a request that says it's GPTBot or Googlebot. To let real AI agents in while keeping impostors out, you need to check where a request actually came from. There are three ways to do it, and the right one depends on the operator.

Why user agents can't be trusted

Every request names its sender in the User-Agent header, and that name is a promise, not a proof. Scrapers routinely borrow the names of well-known crawlers, because many sites let those crawlers through without a challenge. If your bot rules trust the name alone, you're letting in anyone who copies it. If they distrust it, you're blocking real agents along with the fakes.

Verification resolves that. It answers one question: did this request really come from the company it names?

Method 1: published IP ranges

Several operators publish the IP addresses their crawlers use, as machine-readable JSON files. If a request claims to be one of their agents and comes from an address in the list, it's real.

OperatorAgents covered
GoogleGooglebot, special-case crawlers and user-triggered fetchers (separate lists)
MicrosoftBingbot
AppleApplebot
OpenAIGPTBot, OAI-SearchBot and ChatGPT-User (one list each)
PerplexityPerplexityBot and Perplexity-User (one list each)

Each operator links its list from its crawler documentation. A few practical points:

  • Refresh the lists often. Ranges change. Fetch them at least daily, and don't treat a request as fake because it's missing from a list that's days old.
  • Match the list to the agent. An OpenAI address in the GPTBot list doesn't verify a request claiming to be ChatGPT-User.
  • Handle IPv6. Several lists include IPv6 ranges.

Method 2: forward-confirmed reverse DNS

Some operators instead promise that their crawlers' IP addresses resolve to a host name on their own domain. The check has two steps, because reverse DNS on its own can be faked by whoever controls the IP address:

  1. Reverse lookup: look up the host name for the request's IP address, and check it ends in the operator's domain.
  2. Forward lookup: look up that host name's IP addresses, and check the original IP is among them.

Google's own example, from a terminal:

$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.

$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1

The name ends in googlebot.com and points back to the same address, so the request is from Google. Domains that operators document for this include:

AgentHost name ends in
Googlebot and other Google crawlersgooglebot.com, google.com or googleusercontent.com
Bingbotsearch.msn.com
Applebotapplebot.apple.com
Amazonbotcrawl.amazonbot.amazon

DNS lookups are slow compared with a page request, so cache the result for each IP address rather than checking on every hit.

Method 3: signed requests (Web Bot Auth)

The newest method doesn't depend on IP addresses at all. With Web Bot Auth, the agent signs each request with a private key and publishes the matching public key on its own domain. Your CDN or server checks the signature, and a valid one proves the sender, whatever network it came from.

That makes it the best fit for browser agents running in the cloud, whose addresses change constantly. It's still early, so check whether your CDN supports it. We explain how it works, and why we sign all of our own agents' requests, in Why we sign every request our agents send.

When there's no proof at all

Not every operator publishes IP ranges, a reverse DNS domain, or signing keys for every agent. For those, you can't prove a request is real or fake; you can only weigh the evidence, such as how it behaves and how fast it requests pages.

Treat "can't tell" as its own answer, not as "fake". Blocking everything you can't verify will block legitimate agents from operators that simply haven't published proof yet.

Putting it together

  1. Identify: match the user agent to a known agent and its operator.
  2. Verify: use the strongest proof that operator offers: a signature, then published IP ranges, then reverse DNS.
  3. Decide: let verified agents you want through; challenge or block requests that claim a name but fail verification; and handle "can't tell" with ordinary rate limits rather than a hard block.

Your CDN may already do this. Many CDNs and bot-management products maintain lists of verified bots. Check that the AI assistants and AI search agents you want are in the allowed categories, and that "AI crawler" blocking rules aren't catching them by name.

Ghost Agent Labs does it for you. It identifies every agent in your server or edge traffic, checks it against published IP ranges and reverse DNS, and flags impostors, so you can see which agents are real before you decide what to allow. Missing or stale data is reported as "can't tell", never as "spoofed". Start free.

← All articles Test your site with AgentScore →

Can AI agents use your site?

Get your free AgentScore in under a minute. No sign-up needed.