# How to tell if an AI crawler is real

> Anyone can claim to be GPTBot or Googlebot. How to verify AI agents with published IP ranges, reverse DNS and signed requests, and what to do when there's no proof.

Published October 8, 2026 by Ghost Agent Labs · Access · https://ghostagentlab.com/articles/verify-ai-crawlers/

## Key takeaways

- Anyone can send a request claiming to be GPTBot or Googlebot, so trusting the name alone lets impostors straight in.
- Treat requests you can't verify as unknown rather than fake, because blocking all of them can shut out legitimate agents.
- Ask whether your CDN already verifies bots, and make sure the AI assistants and AI search agents you want are allowed.

Anyone can send a request that says it's GPTBot or Googlebot. To let real AI agents in while keeping impostors out, you need to check where a request actually came from. There are three ways to do it, and the right one depends on the operator.

## Why user agents can't be trusted

Every request names its sender in the `User-Agent` header, and that name is a promise, not a proof. Scrapers routinely borrow the names of well-known crawlers, because many sites let those crawlers through without a challenge. If your bot rules trust the name alone, you're letting in anyone who copies it. If they distrust it, you're blocking real agents along with the fakes.

Verification resolves that. It answers one question: did this request really come from the company it names?

## Method 1: published IP ranges

Several operators publish the IP addresses their crawlers use, as machine-readable JSON files. If a request claims to be one of their agents and comes from an address in the list, it's real.

| Operator | Agents covered |
| --- | --- |
| Google | Googlebot, special-case crawlers and user-triggered fetchers (separate lists) |
| Microsoft | Bingbot |
| Apple | Applebot |
| OpenAI | GPTBot, OAI-SearchBot and ChatGPT-User (one list each) |
| Perplexity | PerplexityBot and Perplexity-User (one list each) |

Each operator links its list from its crawler documentation. A few practical points:

- **Refresh the lists often.** Ranges change. Fetch them at least daily, and don't treat a request as fake because it's missing from a list that's days old.
- **Match the list to the agent.** An OpenAI address in the GPTBot list doesn't verify a request claiming to be ChatGPT-User.
- **Handle IPv6.** Several lists include IPv6 ranges.

## Method 2: forward-confirmed reverse DNS

Some operators instead promise that their crawlers' IP addresses resolve to a host name on their own domain. The check has two steps, because reverse DNS on its own can be faked by whoever controls the IP address:

1. **Reverse lookup:** look up the host name for the request's IP address, and check it ends in the operator's domain.
2. **Forward lookup:** look up that host name's IP addresses, and check the original IP is among them.

1. A request arrives
   
   User agent says Googlebot, from IP address `66.249.66.1`
2. 1. Reverse lookup
   
   `66.249.66.1` → `crawl-66-249-66-1.googlebot.com`
   
   Ends in googlebot.com
3. 2. Forward lookup
   
   `crawl-66-249-66-1.googlebot.com` → `66.249.66.1`
   
   Points back to the same address
4. Verified
   
   The request is from Google. If either step fails, it isn’t, whatever its user agent says.

*Forward-confirmed reverse DNS, using Google’s own example. The second step matters because whoever controls an IP address can make its reverse lookup say anything.*

Google's own example, from a terminal:

```
$ host 66.249.66.1
1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.

$ host crawl-66-249-66-1.googlebot.com
crawl-66-249-66-1.googlebot.com has address 66.249.66.1
```

The name ends in `googlebot.com` and points back to the same address, so the request is from Google. Domains that operators document for this include:

| Agent | Host name ends in |
| --- | --- |
| Googlebot and other Google crawlers | `googlebot.com`, `google.com` or `googleusercontent.com` |
| Bingbot | `search.msn.com` |
| Applebot | `applebot.apple.com` |
| Amazonbot | `crawl.amazonbot.amazon` |

DNS lookups are slow compared with a page request, so cache the result for each IP address rather than checking on every hit.

## Method 3: signed requests (Web Bot Auth)

The newest method doesn't depend on IP addresses at all. With [Web Bot Auth](https://blog.cloudflare.com/web-bot-auth/), the agent signs each request with a private key and publishes the matching public key on its own domain. Your CDN or server checks the signature, and a valid one proves the sender, whatever network it came from.

That makes it the best fit for browser agents running in the cloud, whose addresses change constantly. It's still early, so check whether your CDN supports it. We explain how it works, and why we sign all of our own agents' requests, in [Why we sign every request our agents send](https://ghostagentlab.com/blog/signed-requests/).

## When there's no proof at all

Not every operator publishes IP ranges, a reverse DNS domain, or signing keys for every agent. For those, you can't prove a request is real or fake; you can only weigh the evidence, such as how it behaves and how fast it requests pages.

Treat "can't tell" as its own answer, not as "fake". Blocking everything you can't verify will block legitimate agents from operators that simply haven't published proof yet.

## Putting it together

1. **Identify:** match the user agent to a known agent and its operator.
2. **Verify:** use the strongest proof that operator offers: a signature, then published IP ranges, then reverse DNS.
3. **Decide:** let verified agents you want through; challenge or block requests that claim a name but fail verification; and handle "can't tell" with ordinary rate limits rather than a hard block.

> **Your CDN may already do this.** Many CDNs and bot-management products maintain lists of verified bots. Check that the AI assistants and AI search agents you want are in the allowed categories, and that "AI crawler" blocking rules aren't catching them by name.
>
> **Ghost Agent Labs does it for you.** It identifies every agent in your server or edge traffic, checks it against published IP ranges and reverse DNS, and flags impostors, so you can see which agents are real before you decide what to allow. Missing or stale data is reported as "can't tell", never as "spoofed". [Start free](https://app.ghostagentlab.com/signup).

## Sources and further reading

- [Verify requests from Google crawlers and fetchers](https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests) (Google)
- [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots) (OpenAI)
- [About Applebot](https://support.apple.com/en-us/119829) (Apple)
- [HTTP Message Signatures for automated traffic Architecture (Web Bot Auth draft)](https://datatracker.ietf.org/doc/draft-meunier-web-bot-auth-architecture/) (IETF)
