Articles

Setting up your CDN for AI agents

Usually fixed by: Bot protection or CDN admin

Your CDN and web application firewall (WAF) decide, request by request, who gets to see your site. Most were tuned before AI agents were a real source of customers. This guide walks through the settings that matter, in the order your CDN applies them, so verified AI agents get your pages and the bots you don't want still get stopped.

If you're still working out why agents get caught in the first place, start with Bot protection and CAPTCHAs: stop blocking the agents you want. This article is the practical follow-on: how to set the CDN up, and how to check it stays that way.

Start with a decision, not a setting

Before anyone changes a rule, agree which kinds of agent you want. Most businesses land somewhere like this:

Kind of agentWhat it doesTypical choice
AI assistants (for example ChatGPT-User, Claude-User, Perplexity-User)Fetch a page because a person asked a question about it right nowAllow
AI search crawlers (for example OAI-SearchBot)Index pages so they can appear in AI search answersAllow
Search engines (Googlebot, Bingbot)Classic search indexing, which also feeds some AI featuresAllow
AI training crawlers (for example GPTBot, ClaudeBot)Collect content to train modelsA business choice; blocking is common
Unverified bots claiming any of the names aboveUnknownChallenge, rate limit or block

Write the decision down. It's what your CDN admin implements, and it's what you check against later. Our guides to robots.txt for AI agents and whether to block or allow AI crawlers cover the trade-offs.

Use verified bot categories, not user agent strings

Most large CDNs and bot management products keep a list of verified bots: crawlers and agents whose requests they have confirmed really come from the operator they name. Many also group those bots into categories, so you can write one rule for "AI assistants" instead of maintaining a list of names yourself.

Build your rules on that verification, not on the User-Agent header. Anyone can send a request that says it's ChatGPT-User. A rule that skips bot checks for that string is an open door for scrapers. Our guide to telling real AI crawlers from fakes explains the methods CDNs use behind the scenes: published IP ranges, reverse DNS and signatures.

Check the category names in your own console. Vendors name and split these categories differently, and they change them as new agents appear. Confirm which category each agent you care about sits in before you write a rule that depends on it.

Turn on signature verification where you can

IP-based verification struggles with agents that run in the cloud, because their addresses change and are often shared. Web Bot Auth fixes that by having the agent sign each request with a private key, and publish the matching public key on its own domain. It builds on HTTP Message Signatures (RFC 9421), and is still being standardized, so support varies.

If your CDN can verify these signatures, turn it on and treat a valid signature as verification. If it can't yet, nothing breaks: signed requests are ordinary requests with extra headers. Our post Why we sign every request our agents send explains how it works in practice.

Get the rule order right

CDNs apply rules in a sequence, and the first matching action often wins. Most agent problems we see come from order, not from a missing rule. A sensible order looks like this:

  1. Security rules for everyone. Managed WAF rules that catch attacks such as SQL injection should apply to every request, verified bots included. Verification proves who sent a request, not that the request is safe.
  2. Hard blocks you need for legal or business reasons. For example, blocked training crawlers, or country restrictions you're required to enforce. See Geo-blocking, VPN blocks and AI agents for the country case.
  3. Allow verified agents you want, by category or verified identity. "Allow" here should mean "skip bot challenges and bot scoring", not "skip all security".
  4. Rate limits for everyone, including verified agents, set high enough for normal use. See rate limits for AI agents.
  5. Bot scoring and challenges for whatever is left: unverified automation, impostors and suspicious traffic.

The common mistake is a broad "block AI bots" or "challenge automated traffic" rule placed above the verified-agent allow rule. The allow rule then never runs.

Challenge or block: choose deliberately

Most bot tools offer several actions. For a person they feel very different. For an AI agent, most of them end the same way.

ActionWhat a person seesWhat an AI agent gets
BlockAn error pageAn error page; it tells the user it can't reach you
JavaScript or "managed" challengeA short wait, usually invisibleUsually a dead end: most assistant fetchers don't run scripts
Interactive challenge or CAPTCHAA puzzle or checkboxA dead end: agents can't, and shouldn't, solve them
Rate limit (429 with Retry-After)Rarely seenA clear signal to slow down and try again
Log onlyNothingThe page

So treat a challenge as a block when you're thinking about agents. Keep challenges for traffic you genuinely don't want, and for high-risk actions such as login, account creation and payment. Product, pricing, policy and help pages should open for verified agents without one.

Log enough to see what happened

When an assistant says it "couldn't access" your site, someone has to find out why. Make sure your CDN logs, or the samples it keeps, include for each request:

  • The user agent, and whether the CDN treated the request as a verified bot (and in which category)
  • The action taken (allowed, challenged, blocked, rate limited) and the ID of the rule that took it
  • The response status code and the path
  • The client country and network, if you use location rules

Before you switch a new rule to block, run it in log-only (sometimes called "simulate" or "count") mode for a few days and check what it would have caught. Then review blocked and challenged agent traffic regularly. Our guides to measuring AI agent traffic and finding the errors AI agents hit in your logs show what to look for.

Notes for common vendors

Products change quickly and menus move, so treat these as pointers, and check your vendor's current documentation for the exact names and the plan you're on.

  • Cloudflare maintains a verified bots program and documents bot categories that you can reference in custom rules, including separate groupings for AI crawlers and other AI agents. It also offers settings aimed at blocking AI crawlers; check exactly which agents those cover before turning them on. Cloudflare has published work on Web Bot Auth, so check whether signature verification is available on your plan.
  • Akamai bot management sorts known bots into categories and lets you set an action for each. Find where the AI-related categories sit and which action they currently get.
  • Fastly offers bot management and WAF features with its own way of identifying known bots. Check how verified bots are flagged in its rules, and where that check runs relative to your other rules.
  • AWS WAF has a bot control managed rule group that labels requests, including whether a bot is verified. Rules later in the web ACL can act on those labels, so order and labels together decide what agents get.

If you use a dedicated bot product in front of or behind your CDN, the same questions apply to it. Two layers mean two places an agent can be stopped.

Check it from the outside

A configuration that looks right in the console can still behave differently in practice. Test it as an agent would:

  1. Run AgentScore. It requests your home page as a normal browser and as AI agents including ChatGPT-User, Claude-User, Perplexity-User and OAI-SearchBot, and fails the bot protection check if an agent is blocked or challenged while the browser gets through. Other checks test your product, pricing and cart pages the same way, and look for CAPTCHAs on arrival and at checkout.
  2. Look at your CDN's logs for the scan and confirm which rule acted on each request.
  3. Re-run the scan after every bot or WAF change, and after your vendor announces changes to its bot categories.

One caution: AgentScore's requests under those agent names aren't signed and don't come from those operators' networks, so a correctly configured CDN that only allows verified agents may still challenge them. That's the right behavior. Read a failure alongside your logs: if real, verified agent traffic is getting through, your setup is working. Our post on why our scanner wears other names explains the trade-off.

← All articles Test your site with AgentScore →

Can AI agents use your site?

Get your free AgentScore in under a minute. No sign-up needed.