Articles

Bot protection and CAPTCHAs: stop blocking the agents you want

Usually fixed by: Bot protection or CDN admin · Typical effort: hours

Bot protection exists to keep out scrapers, credential stuffers and fraud. But the same rules often stop the AI assistants trying to answer a customer's question about your products. This guide shows how to tell whether that's happening, and how to tune your defenses so the agents you want get through and the bots you don't still get stopped.

Why good agents get caught

Bot protection, whether it's a CDN feature, a web application firewall (WAF) or a dedicated product, looks for signs that a visitor isn't a person. AI agents show many of those signs, for entirely legitimate reasons:

  • They say they're bots. Honest agents announce themselves in their user agent, and a rule that blocks "bots" or "AI crawlers" by name catches them first.
  • They don't run JavaScript. Many challenges work by running a script in the visitor's browser. Assistant fetchers and AI crawlers usually read only the HTML, so they never pass.
  • They come from data centers. Agents run in the cloud, and cloud IP addresses score as higher risk than home broadband.
  • Browser agents look automated. Even agents that drive a real browser move and type differently from people, and fingerprinting tools notice.
  • One-click "block AI" settings. Several providers offer a single switch to block AI bots. Depending on how it's set up, it can block assistants and AI search along with training crawlers.

Signs it's happening to you

  • AI assistants say they "can't access" your site, or describe it from out-of-date information.
  • Your CDN or firewall logs show 403 (forbidden), 429 (too many requests) or challenge responses for agents like ChatGPT-User, Claude-User or Perplexity-User.
  • You appear in AI search results much less often than competitors with weaker SEO.

The quickest test is AgentScore. It requests your home page as a normal browser and as five AI agents, including ChatGPT-User, Claude-User, Perplexity-User and OAI-SearchBot, and compares what comes back. It fails the check if an agent is blocked or challenged while the browser gets through, and warns if an agent gets less than half the content the browser does. Separate checks look for CAPTCHAs on your home page, and test whether agents can open your product, pricing and cart pages.

How to fix it

1. Allow verified agents, not names

The safe way to let agents in is by verified identity. Most bot management products keep a list of verified bots, checked against the operator's published IP ranges, reverse DNS or signatures, and let you allow categories of them. Allow the categories for AI assistants and AI search, and keep blocking requests that only claim those names. Our guide to telling real AI crawlers from fakes explains how verification works.

Never allow by user agent alone. A rule like "if the user agent contains GPTBot, skip all checks" is the first thing scrapers exploit.

2. Separate training from assistants

If you've turned on a "block AI bots" setting, check exactly which agents it covers. If your goal is to opt out of AI training, block only training crawlers, and leave assistant fetchers and AI search crawlers allowed. See robots.txt for AI agents for which names are which.

3. Protect actions, not pages

Abuse mostly targets actions: logging in, creating accounts, applying discount codes, checking gift card balances, submitting payment. It rarely needs protecting against on pages that only display information. Concentrate strict rules and challenges on:

  • Login, sign-up and password reset forms
  • Checkout submission and payment
  • Gift card, coupon and stock-check endpoints that attackers hammer
  • Search and APIs, with rate limits rather than outright blocks

Keep product, category, pricing, policy and help pages as open as you safely can. Those are the pages agents need to answer questions about you.

4. Use rate limits instead of blocks

A real assistant fetcher makes a handful of requests because a person asked about you. A scraper makes thousands. A sensible rate limit per IP address or verified agent stops the scraper without punishing the assistant, where a blanket block stops both.

5. Replace puzzle CAPTCHAs on key pages

AI agents can't solve CAPTCHAs, and aren't meant to. If you need a challenge, use invisible or risk-based challenges that only escalate when traffic looks abusive, and keep them off the pages agents need to read. Never put a CAPTCHA in front of ordinary content on arrival.

6. Don't serve agents a different page

Some setups don't block agents outright, but quietly serve them a stripped-down page, a placeholder, or an "enable JavaScript" message. To the agent this is as bad as a block, and it's harder to notice. Agents should get the same content a browser does.

A quick checklist

  1. Run AgentScore and look at the bot protection, CAPTCHA and key pages checks.
  2. In your CDN or bot management settings, find the verified bots or AI categories, and confirm AI assistants and AI search are allowed.
  3. Check that any "block AI" setting covers only the agents you mean to block.
  4. Search your firewall rules for user agent matches and make sure none allow access by name alone.
  5. Move strict challenges to login, sign-up and checkout submission, and use rate limits elsewhere.
  6. Re-run AgentScore to confirm the fix.

To keep an eye on this over time, watch how often agents are blocked in your traffic. See how to measure AI agent traffic.

← All articles Test your site with AgentScore →

Can AI agents use your site?

Get your free AgentScore in under a minute. No sign-up needed.