Finding the errors AI agents hit, in your server and CDN logs
Usually fixed by: Bot protection or CDN admin
When an AI agent can't get into your site, nobody tells you. There's no complaint and no support ticket. The assistant just answers with someone else's products. The evidence is in your server and CDN logs. This guide shows which errors to look for, how to find them with a few queries, and how to separate real agents from impostors before you change anything.
It builds on how to measure AI agent traffic, which covers getting hold of your logs in the first place.
The errors that matter
| What you see | What it usually means for an agent | Who usually fixes it |
|---|---|---|
| 403 Forbidden | Bot protection, a firewall rule or a country block turned it away | Bot protection or CDN admin |
| 401 Unauthorized | The page needs a login | E-commerce or product team |
| 429 Too Many Requests | It hit a rate limit | CDN admin or developers |
| Challenge page | A "checking your browser" page or CAPTCHA it can't solve. Often a 403 or 503, sometimes a 200 | Bot protection or CDN admin |
| 500, 502, 503 | Your server failed. Agents can reach pages people rarely visit, which may not be cached | Developers or hosting |
| 504, slow responses, 499 | The page took too long. In nginx, 499 means the client gave up and closed the connection first | Developers or hosting |
| 404 Not Found | An old or guessed URL. AI models often remember URLs you've since changed | Content or SEO team |
Challenge pages are the easiest to miss, because some bot protection serves them with a 200 status, so they look like successes. Look for your CDN's security or bot action field instead. Cloudflare, for example, adds a cf-mitigated: challenge header to challenge responses. A response that's much smaller than the real page usually is one too.
What your logs need
Most logs already have the basics. Check yours include:
- Time, client IP address and user agent. The IP address is what lets you verify the agent later.
- Path and status code. Query strings can be dropped; you rarely need them, and they can hold personal data.
- Response size and response time. These catch challenge pages and timeouts.
- The CDN's security action, if it has one: allowed, blocked, challenged, rate limited.
On nginx, a log format like this captures them:
log_format agents '$time_iso8601 $remote_addr "$request_method $uri" '
'$status $body_bytes_sent $request_time "$http_user_agent"';
If you only have the standard "combined" access log, this one-liner counts status codes for three AI assistants:
grep -E 'ChatGPT-User|Claude-User|Perplexity-User' access.log \
| awk '{print $9}' | sort | uniq -c | sort -rn
Example queries
Once logs are in a database or log tool, a few queries answer most questions. The examples below are illustrative, written for PostgreSQL against a table called requests with columns ts, client_ip, user_agent, path, status, bytes, duration_ms and security_action. Rename them to match your own log store; the same ideas work in BigQuery, Athena or a log search tool.
Error rates by agent, last 7 days
SELECT
CASE
WHEN user_agent ILIKE '%ChatGPT-User%' THEN 'ChatGPT-User'
WHEN user_agent ILIKE '%OAI-SearchBot%' THEN 'OAI-SearchBot'
WHEN user_agent ILIKE '%Claude-User%' THEN 'Claude-User'
WHEN user_agent ILIKE '%Perplexity-User%' THEN 'Perplexity-User'
WHEN user_agent ILIKE '%PerplexityBot%' THEN 'PerplexityBot'
WHEN user_agent ILIKE '%Googlebot%' THEN 'Googlebot'
ELSE 'everything else'
END AS agent,
COUNT(*) AS requests,
COUNT(*) FILTER (WHERE status IN (401, 403)) AS blocked,
COUNT(*) FILTER (WHERE status = 429) AS rate_limited,
COUNT(*) FILTER (WHERE status = 404) AS not_found,
COUNT(*) FILTER (WHERE status >= 500) AS server_errors,
ROUND(100.0 * COUNT(*) FILTER (WHERE status >= 400) / COUNT(*), 1) AS error_pct
FROM requests
WHERE ts >= now() - interval '7 days'
GROUP BY 1
ORDER BY requests DESC;
"Everything else" mixes people with other bots, many of them scanners probing for pages that don't exist. For a fair baseline, compare with requests from ordinary browsers: if AI agents get errors far more often than people do, something is turning them away.
Where AI assistants are turned away
SELECT path, status, COUNT(*) AS hits
FROM requests
WHERE ts >= now() - interval '7 days'
AND user_agent ~* '(ChatGPT-User|Claude-User|Perplexity-User)'
AND (status IN (401, 403, 429) OR security_action = 'challenge')
GROUP BY path, status
ORDER BY hits DESC
LIMIT 20;
Challenges and timeouts by day
SELECT date_trunc('day', ts) AS day,
COUNT(*) FILTER (WHERE security_action = 'challenge') AS challenged,
COUNT(*) FILTER (WHERE status IN (499, 504)) AS timed_out,
COUNT(*) FILTER (WHERE duration_ms > 10000) AS over_10s
FROM requests
WHERE user_agent ~* '(ChatGPT-User|Claude-User|Perplexity-User|OAI-SearchBot)'
GROUP BY 1
ORDER BY 1;
A step change on one day usually lines up with a release or a bot protection change. That's your first lead.
Separating verified agents from claimed ones
Everything above trusts the user agent, and a user agent is only a claim. Scrapers often call themselves ChatGPT-User or Googlebot to get past bot protection. When your firewall blocks them, that 403 is working as intended. So before you loosen any rule, split agent errors into three groups:
- Verified. The request came from the operator's published IP ranges, passed forward-confirmed reverse DNS, or carried a valid signature (Web Bot Auth, still an emerging standard). Errors here are real problems to fix.
- Spoofed. It used a known agent's name but failed the operator's check. Blocking these is correct.
- Unverifiable. The operator doesn't publish a way to check. Decide case by case.
Operators such as OpenAI publish their agents' IP ranges (see OpenAI's crawler documentation). Load them into a table, refresh it regularly, and join:
-- agent_ranges(agent text, cidr cidr), loaded from each operator's published list
SELECT r.agent IS NOT NULL AS verified, l.status, COUNT(*) AS hits
FROM requests l
LEFT JOIN agent_ranges r
ON l.client_ip::inet << r.cidr
AND l.user_agent ILIKE '%' || r.agent || '%'
WHERE l.user_agent ~* '(ChatGPT-User|OAI-SearchBot)'
AND l.ts >= now() - interval '7 days'
GROUP BY 1, 2
ORDER BY 1 DESC, hits DESC;
For the methods in detail, including reverse DNS and signed requests, see how to tell if an AI crawler is real.
Mind the privacy. Logs hold IP addresses, which count as personal data for human visitors. Filter to agent user agents before you keep detailed records, hash or drop IP addresses for everyone else, and keep a retention period that matches your privacy policy.
From errors to fixes
- Verified assistants get 403s on product or policy pages: allow verified bots in your bot protection. See bot protection and CAPTCHAs and setting up your CDN for AI agents.
- Repeated 429s: give verified agents their own limits, and send
Retry-After. See rate limits for AI agents. - Blocks only from certain countries or networks: see geo-blocking, VPN blocks and AI agents.
- Challenges at the cart or checkout: these cost orders directly. AgentScore checks for them, and the checkout guide covers the fix.
- The same 404s week after week: redirect old URLs to their current pages.
- Server errors and timeouts: cache the pages agents request most, and check them in your monitoring like any other.
If writing queries isn't your team's idea of a good week, Ghost Agent Labs does this from your Cloudflare, Vercel or JSON logs. The Agent traffic page shows what agents got back, each agent's page shows its status codes, the Verification & spoofing page splits verified from spoofed, and Alerts tell you when a major AI assistant starts being blocked. Either way, check after every bot protection or CDN change, and at least monthly as part of your agent readiness KPIs.