Articles

Finding the errors AI agents hit, in your server and CDN logs

Usually fixed by: Bot protection or CDN admin

When an AI agent can't get into your site, nobody tells you. There's no complaint and no support ticket. The assistant just answers with someone else's products. The evidence is in your server and CDN logs. This guide shows which errors to look for, how to find them with a few queries, and how to separate real agents from impostors before you change anything.

It builds on how to measure AI agent traffic, which covers getting hold of your logs in the first place.

The errors that matter

What you seeWhat it usually means for an agentWho usually fixes it
403 ForbiddenBot protection, a firewall rule or a country block turned it awayBot protection or CDN admin
401 UnauthorizedThe page needs a loginE-commerce or product team
429 Too Many RequestsIt hit a rate limitCDN admin or developers
Challenge pageA "checking your browser" page or CAPTCHA it can't solve. Often a 403 or 503, sometimes a 200Bot protection or CDN admin
500, 502, 503Your server failed. Agents can reach pages people rarely visit, which may not be cachedDevelopers or hosting
504, slow responses, 499The page took too long. In nginx, 499 means the client gave up and closed the connection firstDevelopers or hosting
404 Not FoundAn old or guessed URL. AI models often remember URLs you've since changedContent or SEO team

Challenge pages are the easiest to miss, because some bot protection serves them with a 200 status, so they look like successes. Look for your CDN's security or bot action field instead. Cloudflare, for example, adds a cf-mitigated: challenge header to challenge responses. A response that's much smaller than the real page usually is one too.

What your logs need

Most logs already have the basics. Check yours include:

  • Time, client IP address and user agent. The IP address is what lets you verify the agent later.
  • Path and status code. Query strings can be dropped; you rarely need them, and they can hold personal data.
  • Response size and response time. These catch challenge pages and timeouts.
  • The CDN's security action, if it has one: allowed, blocked, challenged, rate limited.

On nginx, a log format like this captures them:

log_format agents '$time_iso8601 $remote_addr "$request_method $uri" '
                  '$status $body_bytes_sent $request_time "$http_user_agent"';

If you only have the standard "combined" access log, this one-liner counts status codes for three AI assistants:

grep -E 'ChatGPT-User|Claude-User|Perplexity-User' access.log \
  | awk '{print $9}' | sort | uniq -c | sort -rn

Example queries

Once logs are in a database or log tool, a few queries answer most questions. The examples below are illustrative, written for PostgreSQL against a table called requests with columns ts, client_ip, user_agent, path, status, bytes, duration_ms and security_action. Rename them to match your own log store; the same ideas work in BigQuery, Athena or a log search tool.

Error rates by agent, last 7 days

SELECT
  CASE
    WHEN user_agent ILIKE '%ChatGPT-User%'    THEN 'ChatGPT-User'
    WHEN user_agent ILIKE '%OAI-SearchBot%'   THEN 'OAI-SearchBot'
    WHEN user_agent ILIKE '%Claude-User%'     THEN 'Claude-User'
    WHEN user_agent ILIKE '%Perplexity-User%' THEN 'Perplexity-User'
    WHEN user_agent ILIKE '%PerplexityBot%'   THEN 'PerplexityBot'
    WHEN user_agent ILIKE '%Googlebot%'       THEN 'Googlebot'
    ELSE 'everything else'
  END AS agent,
  COUNT(*) AS requests,
  COUNT(*) FILTER (WHERE status IN (401, 403)) AS blocked,
  COUNT(*) FILTER (WHERE status = 429)         AS rate_limited,
  COUNT(*) FILTER (WHERE status = 404)         AS not_found,
  COUNT(*) FILTER (WHERE status >= 500)        AS server_errors,
  ROUND(100.0 * COUNT(*) FILTER (WHERE status >= 400) / COUNT(*), 1) AS error_pct
FROM requests
WHERE ts >= now() - interval '7 days'
GROUP BY 1
ORDER BY requests DESC;

"Everything else" mixes people with other bots, many of them scanners probing for pages that don't exist. For a fair baseline, compare with requests from ordinary browsers: if AI agents get errors far more often than people do, something is turning them away.

Where AI assistants are turned away

SELECT path, status, COUNT(*) AS hits
FROM requests
WHERE ts >= now() - interval '7 days'
  AND user_agent ~* '(ChatGPT-User|Claude-User|Perplexity-User)'
  AND (status IN (401, 403, 429) OR security_action = 'challenge')
GROUP BY path, status
ORDER BY hits DESC
LIMIT 20;

Challenges and timeouts by day

SELECT date_trunc('day', ts) AS day,
  COUNT(*) FILTER (WHERE security_action = 'challenge')   AS challenged,
  COUNT(*) FILTER (WHERE status IN (499, 504))            AS timed_out,
  COUNT(*) FILTER (WHERE duration_ms > 10000)             AS over_10s
FROM requests
WHERE user_agent ~* '(ChatGPT-User|Claude-User|Perplexity-User|OAI-SearchBot)'
GROUP BY 1
ORDER BY 1;

A step change on one day usually lines up with a release or a bot protection change. That's your first lead.

Separating verified agents from claimed ones

Everything above trusts the user agent, and a user agent is only a claim. Scrapers often call themselves ChatGPT-User or Googlebot to get past bot protection. When your firewall blocks them, that 403 is working as intended. So before you loosen any rule, split agent errors into three groups:

  • Verified. The request came from the operator's published IP ranges, passed forward-confirmed reverse DNS, or carried a valid signature (Web Bot Auth, still an emerging standard). Errors here are real problems to fix.
  • Spoofed. It used a known agent's name but failed the operator's check. Blocking these is correct.
  • Unverifiable. The operator doesn't publish a way to check. Decide case by case.

Operators such as OpenAI publish their agents' IP ranges (see OpenAI's crawler documentation). Load them into a table, refresh it regularly, and join:

-- agent_ranges(agent text, cidr cidr), loaded from each operator's published list
SELECT r.agent IS NOT NULL AS verified, l.status, COUNT(*) AS hits
FROM requests l
LEFT JOIN agent_ranges r
  ON l.client_ip::inet << r.cidr
 AND l.user_agent ILIKE '%' || r.agent || '%'
WHERE l.user_agent ~* '(ChatGPT-User|OAI-SearchBot)'
  AND l.ts >= now() - interval '7 days'
GROUP BY 1, 2
ORDER BY 1 DESC, hits DESC;

For the methods in detail, including reverse DNS and signed requests, see how to tell if an AI crawler is real.

Mind the privacy. Logs hold IP addresses, which count as personal data for human visitors. Filter to agent user agents before you keep detailed records, hash or drop IP addresses for everyone else, and keep a retention period that matches your privacy policy.

From errors to fixes

If writing queries isn't your team's idea of a good week, Ghost Agent Labs does this from your Cloudflare, Vercel or JSON logs. The Agent traffic page shows what agents got back, each agent's page shows its status codes, the Verification & spoofing page splits verified from spoofed, and Alerts tell you when a major AI assistant starts being blocked. Either way, check after every bot protection or CDN change, and at least monthly as part of your agent readiness KPIs.

← All articles Test your site with AgentScore →

See which AI agents visit your site

Ghost Agent Labs shows every crawler, AI assistant and browser agent that visits, from your CDN or server logs, and which ones are real. Included in every plan, even Free.