{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Ghost Agent Labs",
  "home_page_url": "https://ghostagentlab.com/",
  "feed_url": "https://ghostagentlab.com/feed.json",
  "description": "Blog posts and guides on making websites work for AI agents.",
  "icon": "https://ghostagentlab.com/assets/feed-icon-512.png",
  "favicon": "https://ghostagentlab.com/assets/feed-icon-144.png",
  "language": "en",
  "authors": [
    {
      "name": "Ghost Agent Labs",
      "url": "https://ghostagentlab.com/"
    }
  ],
  "items": [
    {
      "id": "https://ghostagentlab.com/articles/agent-readiness-glossary/",
      "url": "https://ghostagentlab.com/articles/agent-readiness-glossary/",
      "title": "Agent readiness glossary: the terms you'll hear, in plain English",
      "summary": "Plain-English definitions of the terms behind agent readiness, from assistant fetchers and Web Bot Auth to structured data, the accessibility tree, MCP and AgentScore.",
      "content_html": "<p>Agent readiness comes with a lot of new words, and some old ones used in new ways. This glossary explains the terms you'll meet in our guides, in AgentScore reports and in conversations with your developers and vendors, in plain English and with what each one means for your business.</p>\n          <h2>The agents</h2>\n          <p><strong>AI agent.</strong> Software that uses an AI model to do something on the web for a person or a company: read a page, answer a question, compare products or complete a task. In our guides, “agent” covers everything from crawlers to agents that shop.</p>\n          <p><strong>Training crawler.</strong> A bot that collects pages to train AI models, such as <code>GPTBot</code> or <code>ClaudeBot</code>. Blocking it doesn't stop AI assistants visiting for a person.</p>\n          <p><strong>AI search crawler.</strong> A bot that indexes pages so AI search and assistants can cite and link to them, such as <code>OAI-SearchBot</code> or <code>PerplexityBot</code>.</p>\n          <p><strong>Assistant fetcher.</strong> An AI assistant fetching a page in real time because a person asked it something, such as <code>ChatGPT-User</code> or <code>Claude-User</code>. Each visit stands in for a customer.</p>\n          <p><strong>Browser agent.</strong> An agent that opens your site in a real browser and clicks, types and scrolls like a person, to book, sign up or buy. Also called a computer-use agent or autonomous agent.</p>\n          <p><strong>User agent.</strong> The name a visitor sends with every request to say what it is, such as a browser version or <code>GPTBot</code>. Easy to fake, so it shouldn't be trusted on its own.</p>\n          <p>For how these differ and how to treat each, see <a href=\"https://ghostagentlab.com/articles/kinds-of-ai-agents/\">the AI agents that visit your website</a>.</p>\n          <h2>Getting in: access</h2>\n          <p><strong>robots.txt.</strong> A file at the root of your site that tells crawlers and agents which pages they may visit. Well-behaved agents follow it; scrapers ignore it. See <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>.</p>\n          <p><strong>Bot protection.</strong> Services such as Cloudflare, Akamai or DataDome that block or challenge automated visitors. Set too strictly, they turn away AI assistants as well as scrapers. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a>.</p>\n          <p><strong>CDN.</strong> Content delivery network: the service that sits in front of your site and delivers it quickly around the world. It's often where bot rules live, and where agent traffic is recorded. See <a href=\"https://ghostagentlab.com/articles/cdn-settings-ai-agents/\">setting up your CDN for AI agents</a>.</p>\n          <p><strong>CAPTCHA or challenge.</strong> A test that tries to tell people from bots, such as picking out images or waiting on a “checking your browser” page. Agents can't pass them, so one on arrival or at checkout ends the visit.</p>\n          <p><strong>Rate limit.</strong> A cap on how many requests a visitor can make in a given time. Fair limits stop scrapers without turning away agents. See <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a>.</p>\n          <p><strong>Verification.</strong> Checking that a visitor claiming to be a known agent really is, using the proof its operator publishes: IP ranges, reverse DNS or a signature. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</p>\n          <p><strong>Impostor (spoofed agent).</strong> A request that uses a trusted agent's name but fails verification. Usually a scraper trying to get past bot protection.</p>\n          <p><strong>Forward-confirmed reverse DNS.</strong> A verification check that an IP address's hostname belongs to the operator, and that the hostname points back to the same address.</p>\n          <p><strong>Web Bot Auth.</strong> A new standard in which agents sign their requests cryptographically, so sites can prove who sent them. Our own scanner and Ghost Agents sign theirs. See <a href=\"https://ghostagentlab.com/blog/signed-requests/\">why we sign every request</a>.</p>\n          <p><strong>llms.txt.</strong> A short Markdown file at <code>/llms.txt</code> that tells AI what your business does and links to your key pages. An emerging standard. See <a href=\"https://ghostagentlab.com/articles/llms-txt/\">how to write an llms.txt file</a>.</p>\n          <p><strong>Guest checkout.</strong> Letting shoppers buy without creating an account. An agent buying for someone can't easily sign up for them. See <a href=\"https://ghostagentlab.com/articles/guest-checkout-ai-agents/\">guest checkout</a>.</p>\n          <h2>Understanding pages: readability</h2>\n          <p><strong>Raw HTML.</strong> The page exactly as your server sends it, before any JavaScript runs. It's all that many crawlers and assistant fetchers ever see. To look at it, use your browser's “View page source”, not the inspector.</p>\n          <p><strong>Server-side rendering (SSR) and static generation.</strong> Ways of building pages so the content is already in the raw HTML, rather than added afterwards by JavaScript. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see JavaScript-only content</a>.</p>\n          <p><strong>Structured data (schema.org, JSON-LD).</strong> Machine-readable facts in your page code, such as a product's name, price and stock, written in the shared schema.org vocabulary, usually as JSON-LD. Agents can read it without guessing. See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data AI shopping agents can read</a>.</p>\n          <p><strong>XML sitemap.</strong> A file listing the pages you want found, so crawlers and agents don't depend on following links. See <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">XML sitemaps for AI agents</a>.</p>\n          <p><strong>Canonical URL.</strong> The one address you name as the main version of a page that can be reached in several ways. It stops agents treating duplicates as different products. See <a href=\"https://ghostagentlab.com/articles/canonical-urls-ai-agents/\">canonical URLs</a>.</p>\n          <p><strong>Alt text.</strong> A short text description of an image. Agents can't see pictures, so this is how they learn what a product photo shows. See <a href=\"https://ghostagentlab.com/articles/image-alt-text-ai-agents/\">alt text for AI agents</a>.</p>\n          <h2>Finding the way: navigability</h2>\n          <p><strong>Accessibility tree.</strong> The structured view of a page that screen readers and many browser agents use: its headings, links, buttons and form fields, with their names. If something isn't in it, many agents can't use it. See <a href=\"https://ghostagentlab.com/blog/what-an-ai-agent-sees/\">what an AI agent sees</a>.</p>\n          <p><strong>Accessible name.</strong> The name a button, link or field has in the accessibility tree, from its text, label or <code>aria-label</code>. An icon-only button with no name is invisible to agents.</p>\n          <p><strong>Landmarks.</strong> Labels for the main regions of a page, such as header, navigation, main content and footer, that let agents skip straight to what they need.</p>\n          <p><strong>Overlay.</strong> Anything that covers the page, such as a cookie banner, newsletter pop-up or chat window. If an agent can't close it, the journey ends there. See <a href=\"https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/\">cookie banners and pop-ups</a>.</p>\n          <p><strong>WCAG.</strong> The Web Content Accessibility Guidelines, the standard for making sites usable by people with disabilities. Much of it helps agents too. See <a href=\"https://ghostagentlab.com/blog/accessibility-is-agent-readiness/\">accessibility work is agent readiness work</a>.</p>\n          <h2>Getting it done: tasks and commerce</h2>\n          <p><strong>Journey.</strong> A task a visitor sets out to complete on your site, such as finding a product and reaching checkout, or booking a call. The journeys that make you money are the ones to test first.</p>\n          <p><strong>Task completion.</strong> Whether an agent can actually finish a journey. It's the AgentScore category worth the most points, and the only way to measure it is to try. See <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">how to test your key journeys</a>.</p>\n          <p><strong>MCP (Model Context Protocol).</strong> A standard way to give AI agents direct tools for your site, such as search, product details and cart, instead of making them click through pages.</p>\n          <p><strong>WebMCP.</strong> A proposal for offering those same tools from inside your web pages, so a browser agent can call them while it's on your site. See <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP and WebMCP</a>.</p>\n          <p><strong>Agentic commerce.</strong> Purchases that an AI agent makes for a shopper. Protocols such as the Agentic Commerce Protocol, the Universal Commerce Protocol and AP2 aim to make that safe for buyer, store and payment provider. See <a href=\"https://ghostagentlab.com/articles/agentic-commerce-protocols/\">agentic commerce protocols explained</a>.</p>\n          <h2>Measuring it</h2>\n          <p><strong>Agent readiness.</strong> How well AI agents can get into your site, understand it, find their way around it and complete tasks on it. See <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">what is agent readiness?</a></p>\n          <p><strong>AgentScore.</strong> Our free score out of 100 for agent readiness, across Access (25 points), Readability (25), Navigability (20) and Task completion (30). The free scan marks Task completion “not tested” and scores the other three. See <a href=\"https://ghostagentlab.com/blog/how-agentscore-works/\">how AgentScore works</a>.</p>\n          <p><strong>Check.</strong> One thing AgentScore tests, such as “Buttons and links have names”, reported as a pass, warning or fail, with who usually fixes it.</p>\n          <p><strong>Ghost Agent.</strong> A real AI agent we send through a journey on your site, on a schedule, to see whether agents can complete it. You get a step-by-step replay of each run, and an alert when a journey breaks. See <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a>.</p>\n          <p><strong>Run and pass rate.</strong> A run is one attempt by one agent at a journey. AI agents don't behave the same way every time, so journeys are tried several times, and the pass rate is the share of runs that succeed.</p>\n          <p><strong>Flaky.</strong> A journey where some runs pass and some fail. Agents can do it, but not reliably, which usually points to a slow page, a pop-up that appears sometimes, or an unclear control.</p>\n          <p><strong>Agent traffic.</strong> Visits from AI agents of every kind. Most never run your analytics tag, so it's measured from server or CDN logs. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure agent traffic</a>.</p>\n          <p><strong>AI referral.</strong> A person who arrives at your site by clicking a link in an AI assistant's answer. Unlike agent visits, these do show up in analytics. See <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">tracking visits and sales from AI assistants</a>.</p>\n          <p><strong>Crawls per AI referral.</strong> How many pages an AI company's agents fetched for every visitor its assistant sent you. A low number means you get traffic back for what they read.</p>\n          <p>To see where your own site stands on each of these, run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>. It's free and takes under a minute.</p>",
      "date_published": "2026-10-09T13:26:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Start here"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/kinds-of-ai-agents/",
      "url": "https://ghostagentlab.com/articles/kinds-of-ai-agents/",
      "title": "The AI agents that visit your website, and what each one wants",
      "summary": "Training crawlers, AI search crawlers, assistant fetchers, browser agents and tool-using agents: what each does, what it needs from your site, and how to treat it.",
      "content_html": "<p>“AI bots” isn't one kind of visitor. Some collect pages to train models, some build AI search, some fetch a page because a person just asked about you, and some open your site in a browser and click through it. This guide explains the five kinds, what each one needs from your site, and how to treat each.</p>\n          <h2>Why the difference matters</h2>\n          <p>Each kind of agent does a different job for a different company, and brings you a different kind of value. A training crawler takes your content and may never send anyone back. An assistant fetching your returns policy is answering a customer's question right now. A browser agent on your checkout page is trying to place an order.</p>\n          <p>Treat them all as “bots” and you get one of two bad outcomes: block everything and disappear from AI answers and purchases, or allow everything without knowing what you've agreed to. The rest of this guide gives you the vocabulary to make a separate decision for each.</p>\n          <h2>The five kinds at a glance</h2>\n          <table>\n            <thead><tr><th>Kind</th><th>Why it visits</th><th>How it reads your site</th><th>What it's worth to you</th></tr></thead>\n            <tbody>\n              <tr><td>Training crawler</td><td>To collect pages for training AI models</td><td>Raw HTML, in bulk</td><td>Indirect at best</td></tr>\n              <tr><td>AI search crawler</td><td>To index pages so AI search can cite and link to them</td><td>Mostly raw HTML</td><td>Visibility in AI answers</td></tr>\n              <tr><td>Assistant fetcher</td><td>A person asked an assistant about a specific page or question</td><td>Usually raw HTML, one page at a time</td><td>A customer asking about you right now</td></tr>\n              <tr><td>Browser agent</td><td>A person asked an assistant to do something, like book or buy</td><td>A real browser: renders the page, clicks and types</td><td>A task, often a sale, in progress</td></tr>\n              <tr><td>Tool-using agent</td><td>To call tools your site offers directly, through MCP or an API</td><td>Structured answers from your tools, not pages</td><td>The fastest, most reliable route to a task</td></tr>\n            </tbody>\n          </table>\n          <h2>1. Training crawlers</h2>\n          <p>Training crawlers collect large numbers of pages to train AI models. Examples include <code>GPTBot</code>, <code>ClaudeBot</code> and <code>CCBot</code>. They don't act for a particular person, and a visit from one doesn't mean anyone is looking for you.</p>\n          <ul>\n            <li><strong>What they need:</strong> nothing special. They read your HTML like a search crawler.</li>\n            <li><strong>How to control them:</strong> robots.txt. They use their own names, so you can block them without blocking AI search or assistants. See <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>.</li>\n            <li><strong>The decision:</strong> whether your content may be used for training is a business and legal choice, not a readiness one. AgentScore reports blocked training crawlers but doesn't take points off for them.</li>\n          </ul>\n          <h2>2. AI search crawlers</h2>\n          <p>AI search crawlers, such as <code>OAI-SearchBot</code>, <code>Claude-SearchBot</code> and <code>PerplexityBot</code>, build the index that AI search and assistants draw on when they answer questions and cite sources. They work much like Googlebot does for traditional search.</p>\n          <ul>\n            <li><strong>What they need:</strong> your key content in the HTML your server sends, clear titles and headings, structured data, and a sitemap. Many don't run JavaScript. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see JavaScript-only content</a>.</li>\n            <li><strong>How to control them:</strong> robots.txt, and your CDN or bot protection, which has to let them through.</li>\n            <li><strong>Why allow them:</strong> if they can't read you, AI answers about your category are written from your competitors' pages.</li>\n          </ul>\n          <h2>3. Assistant fetchers</h2>\n          <p>When someone asks ChatGPT, Claude or Perplexity “what's the return window at this store?” or “compare these two products”, the assistant may fetch the pages it needs there and then. Those requests use names like <code>ChatGPT-User</code>, <code>Claude-User</code> and <code>Perplexity-User</code>.</p>\n          <ul>\n            <li><strong>What they need:</strong> fast, readable pages with the facts in plain text: prices, stock, shipping and returns. See <a href=\"https://ghostagentlab.com/articles/policy-pages-ai-assistants/\">shipping, returns and FAQ pages AI assistants can quote</a>.</li>\n            <li><strong>How to control them:</strong> some operators treat these fetches like a person clicking a link and say robots.txt may not apply. Stopping them completely means a rule at your CDN or server.</li>\n            <li><strong>Why allow them:</strong> every one of these visits stands in for a customer asking about you. These are usually the visits you want most.</li>\n          </ul>\n          <h2>4. Browser agents</h2>\n          <p>Browser agents go further: they open your site in a real browser and use it the way a person would, to search, choose a size, fill in a form or reach checkout. Several AI assistants now offer an agent mode that does this on a person's behalf, and businesses are starting to use them too.</p>\n          <ul>\n            <li><strong>What they need:</strong> buttons, links and form fields with clear names, pop-ups they can close, no CAPTCHA in the way, and a checkout that doesn't force an account. They find their way around much as a screen reader does. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a> and <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">how to make checkout work for AI shopping agents</a>.</li>\n            <li><strong>How to recognize them:</strong> often you can't from the name alone, because many use an ordinary browser user agent. Some operators now sign their requests with <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">Web Bot Auth</a> so sites can tell who they are.</li>\n            <li><strong>Why they matter:</strong> they're the agents that complete tasks, so they're where agent readiness turns into revenue, and where a broken button or a bot challenge costs you an order.</li>\n          </ul>\n          <h2>5. Tool-using agents</h2>\n          <p>Instead of reading pages, some agents call tools a site offers them directly: “search products”, “check stock”, “add to cart”. These tools are offered through the Model Context Protocol (MCP), WebMCP in the browser, or agentic commerce protocols for checkout.</p>\n          <ul>\n            <li><strong>What they need:</strong> a set of well-described, safe tools, starting with read-only ones such as search and product details.</li>\n            <li><strong>Why it's worth planning for:</strong> tools are faster and more reliable than clicking through pages, and few sites offer them yet. See <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP and WebMCP</a> and <a href=\"https://ghostagentlab.com/articles/agentic-commerce-protocols/\">agentic commerce protocols explained</a>.</li>\n            <li><strong>Where to start:</strong> most businesses get more from fixing the first four kinds first. Tools build on readable pages and clean product data, rather than replacing them.</li>\n          </ul>\n          <h2>How to see which ones visit you</h2>\n          <p>Your web analytics probably shows almost none of this. Most agents never run the JavaScript tag analytics depends on, and the ones that do are often filtered out as bots. To see them, look at your server or CDN logs, where every request is recorded with its user agent and IP address. Then check that the visitors claiming to be well-known agents really are: names are easy to fake.</p>\n          <ol>\n            <li><strong>Count by kind, not by bot.</strong> Group visits into the five kinds above, so you can see how many customer-driven fetches you get compared with crawls.</li>\n            <li><strong>Verify the big names.</strong> Check claimed agents against the IP ranges or signatures their operators publish. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</li>\n            <li><strong>Watch for errors.</strong> A rise in blocked or failed requests from one kind of agent usually means a rule changed. See <a href=\"https://ghostagentlab.com/articles/agent-errors-in-logs/\">finding the errors AI agents hit</a>.</li>\n          </ol>\n          <p>Ghost Agent Labs does this for you from your CDN or server logs. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure agent traffic</a>.</p>\n          <h2>A sensible starting policy</h2>\n          <div>\n            <ul>\n              <li><strong>Training crawlers:</strong> your choice. Block them by name in robots.txt if you don't want your content used for training.</li>\n              <li><strong>AI search crawlers and assistant fetchers:</strong> allow them, in robots.txt and at your CDN, and make sure your key content is in the HTML.</li>\n              <li><strong>Browser agents:</strong> allow verified ones through bot protection, especially on product, cart and checkout pages, and test that they can finish your key journeys.</li>\n              <li><strong>Tool-using agents:</strong> plan for them once the basics work, starting with read-only tools.</li>\n              <li><strong>Everyone else:</strong> keep blocking unknown and abusive bots, and anything that claims a trusted name but fails verification.</li>\n            </ul>\n          </div>\n          <p>For a fuller decision guide by business type, see <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">should you block AI crawlers?</a> To see how your site treats each kind today, run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>: its robots.txt check reads your rules as 21 AI agents would.</p>",
      "date_published": "2026-10-09T13:25:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Start here"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/what-scanning-600-stores-taught-us/",
      "url": "https://ghostagentlab.com/blog/what-scanning-600-stores-taught-us/",
      "title": "What scanning 600 stores taught us about our own scanner",
      "summary": "Running AgentScore on 600 real stores exposed three bugs in our own scanner, and one run that looked perfect because almost nothing was tested. What we found, fixed and changed.",
      "content_html": "<p>To write <a href=\"https://ghostagentlab.com/blog/agentscore-store-study-2026/\">our study of online stores</a>, we ran AgentScore on 600 real websites in one day. The study taught us a lot about stores. It taught us even more about our own scanner: three bugs that tests on our own fixture sites never caught, and one run whose results looked perfect because almost nothing had been tested. Here's what we found, what we changed, and what it means for your score.</p>\n          <h2>Why run the scanner at scale</h2>\n          <p>Every AgentScore check is tested against small, carefully built example sites. That proves the check does what we meant. It doesn't prove we meant the right thing for the thousands of ways real stores are built. Scanning 600 stores, from many countries, platforms and sizes, was the first time AgentScore met that variety all at once, and the first time we read its results in aggregate, where a check that's wrong for many stores stands out.</p>\n          <h2>Bug 1: AgentScore only spoke English</h2>\n          <p>Several checks work by reading what a button or link says: finding the add-to-cart button, recognising the button that closes a cookie banner, spotting the link to the cart. All of them only knew English words. On a German store, \"In den Warenkorb\" wasn't recognised as an add-to-cart button, and \"Alle akzeptieren\" didn't count as a way to close the banner.</p>\n          <p>In our first run, 61 stores whose pages weren't in English were reported as having no add-to-cart button at all. When we rescanned eight of them with the fix, AgentScore found the button on six. The other two turned out to be a different problem (see bug 3).</p>\n          <p>There was a subtler issue underneath. The usual way to match whole words in code only understands the 26 letters of the English alphabet, so even a correctly translated phrase could fail to match at the edges of words in Russian or with accented letters. Japanese, Chinese and Korean don't put spaces between words at all.</p>\n          <p><strong>What we changed:</strong> the words AgentScore looks for now live in one list covering 19 languages, including German, French, Spanish, Italian, Dutch, the Nordic languages, Polish, Turkish, Russian, Ukrainian, Japanese, Chinese and Korean, and matching works in every alphabet. It also changed which sites AgentScore recognises as stores, because finding the cart link is part of that: the same sample produced 230 stores instead of 219.</p>\n          <h2>Bug 2: Real buttons reported as fake ones</h2>\n          <p>AgentScore checks that your add-to-cart button is a real button, because agents look for buttons and links, not for text that happens to be clickable. To find the button, it looked for the smallest element whose text matched. Many shop themes write their buttons like this:</p>\n          <pre><code>&lt;button type=\"submit\"&gt;\n  &lt;span&gt;Add to cart&lt;/span&gt;\n&lt;/button&gt;</code></pre>\n          <p>The smallest element containing \"Add to cart\" is the <code>&lt;span&gt;</code>, so AgentScore reported \"a plain &lt;span&gt;, not a real button\", on a perfectly good button. That's a false alarm, and an expensive one: in the study's second run, 57% of the stores where we found the button were flagged. After the fix, the true figure is 24%.</p>\n          <p><strong>What we changed:</strong> when the matching text sits inside a real button or link, AgentScore now judges that button or link. A <code>&lt;div&gt;</code> styled to look like a button is still caught, because that really is a problem for agents.</p>\n          <h2>Bug 3: Sold-out products counted against stores</h2>\n          <p>AgentScore checks one product page per store. When that product was sold out, its buy button had been replaced by a disabled \"Sold out\" button, \"売り切れ\" on a Japanese store or \"Agotado\" on a Spanish one, and AgentScore reported that it couldn't find an add-to-cart button. That says nothing about the store's buttons, only about its stock.</p>\n          <p><strong>What we changed:</strong> AgentScore now recognises a sold-out product, from its button in any of the 19 languages or from the product data saying it's out of stock. Then it does what a shopper would: it tries up to two other products. If they're all sold out, the check is marked as not tested, which never costs points. In the study, AgentScore moved past a sold-out product on 4 stores and skipped the check on 5 where everything it tried was sold out.</p>\n          <h2>The run that looked perfect</h2>\n          <p>One of our runs came back looking wonderful. Navigation scored 100%. Not one store had an unnamed button. No pop-ups blocked anything.</p>\n          <p>It was wrong. After an update to how AgentScore opens pages in its browser, the browser on our scanning machine couldn't complete secure connections, and it couldn't load a single page. Every check that needs a real browser was skipped, on 245 of 245 stores. Skipped checks don't cost points, by design, so the scores went up instead of down.</p>\n          <p>We caught it because the numbers were too good to be true, not because anything failed. That's not good enough, so we added a safeguard: our analysis now refuses to produce results if more than 10% of stores couldn't be opened in a browser. We fixed the scanning machine's setup, confirmed one scan by hand, and ran all 600 sites again.</p>\n          <div>\n            <p><strong>Your own scans are protected the same way.</strong> When AgentScore can't open your pages in a browser, your report says so and marks those checks as not tested. It never presents a skipped check as a pass. See <a href=\"https://ghostagentlab.com/blog/reading-your-agentscore-report/\">how to read your AgentScore report</a>.</p>\n          </div>\n          <h2>What this means for your score</h2>\n          <ul>\n            <li><strong>Stores outside English-speaking markets</strong> get a fairer score, and some will see it rise, as buttons, banners and cart links are now recognised in their language.</li>\n            <li><strong>Stores whose buttons wrap their text</strong> in another element no longer get a false \"not a real button\" finding.</li>\n            <li><strong>A sold-out product</strong> no longer costs you points.</li>\n          </ul>\n          <p>All three changes are in AgentScore's current scoring version, 2026.10-v5. If you scanned your site before, <a href=\"https://ghostagentlab.com/agentscore/\">run AgentScore again</a> to see your updated score.</p>\n          <h2>What we took away</h2>\n          <ul>\n            <li><strong>Test against the real world, not only your examples.</strong> Each bug passed every test we had, because our test sites were built the way we expected sites to be built.</li>\n            <li><strong>Be suspicious of good news.</strong> The broken run would have made a better headline than the real one. A result that's surprisingly good deserves the same scrutiny as one that's surprisingly bad.</li>\n            <li><strong>Publish the limits.</strong> The study states what AgentScore can't see, such as checkouts that need items in the cart and requests that don't come from AI companies' verified addresses. Saying so is what makes the rest of the numbers worth trusting.</li>\n          </ul>\n          <p>Curious how AgentScore opens and reads your pages? <a href=\"https://ghostagentlab.com/blog/how-agentscore-renders-pages/\">How AgentScore renders pages</a> explains the browser side, and <a href=\"https://ghostagentlab.com/blog/how-agentscore-works/\">How AgentScore works</a> covers the scoring.</p>",
      "date_published": "2026-10-09T13:24:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Engineering"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/agentscore-store-study-2026/",
      "url": "https://ghostagentlab.com/blog/agentscore-store-study-2026/",
      "title": "We scanned 234 online stores with AgentScore. Here's what trips up AI agents",
      "summary": "Most online stores let AI agents in, but many trip them up once inside: unnamed buttons, missing product data, and prices agents can't see. Results from 234 stores.",
      "content_html": "<p>AI shopping agents are starting to research and buy on people's behalf. So we asked a simple question: if an AI agent visited a typical online store today, how far would it get? We ran AgentScore on 234 online stores to find out. The short answer: most stores let agents in, but many trip them up once they're inside.</p>\n          <div>\n            <div><strong>85</strong><p>Median AgentScore, out of 100</p></div>\n            <div><strong>48%</strong><p>of stores have buttons or links with no name an agent can read</p></div>\n            <div><strong>30%</strong><p>of stores don't put prices in the product page's HTML, where most agents look</p></div>\n          </div>\n          <h2>How we did it</h2>\n          <p>We drew 600 websites at random from the online shops listed in <a href=\"https://www.wikidata.org/\">Wikidata</a>, the public knowledge base, so the sample leans toward established retailers notable enough to have an entry, from many countries. On October 9, 2026 we ran the same automated <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> scan (scoring version 2026.10-v5) on each, and 544 could be scanned; the rest were unreachable, timed out or returned errors. 234 of those sites were identified as stores, with products or a cart, and those are the ones in this report. Another 12 stores showed a bot challenge even to an ordinary browser request from our cloud servers, so we couldn't see them at all and left them out.</p>\n          <p>For each store, AgentScore checks three things:</p>\n          <ul>\n            <li><strong>Access:</strong> whether robots.txt rules, bot protection or CAPTCHAs stop AI assistants such as ChatGPT-User, Claude-User and Perplexity-User from reaching the home page, product pages and checkout.</li>\n            <li><strong>Readability:</strong> whether content, product details and prices are in the page's HTML, or only appear after JavaScript runs.</li>\n            <li><strong>Navigability:</strong> whether buttons, links and form fields have names agents can read, whether pop-ups block the page, and whether the add-to-cart button is a real, uncovered button.</li>\n          </ul>\n          <p>This version doesn't send an AI agent through a live shopping task, so the scores cover whether agents can get in, read and find their way, not whether they complete a purchase. The scan is read-only: it never submits forms, adds to cart or places orders. All scans ran from a single cloud location, so pop-ups and content that vary by country may differ for visitors elsewhere. We report results only in aggregate and don't name individual stores.</p>\n          <h2>The scores</h2>\n          <p>The median store scored 85 out of 100. 65% scored 80 or more, and 1% scored below 50.</p>\n          <table>\n            <thead><tr><th>AgentScore</th><th>Share of stores</th></tr></thead>\n            <tbody><tr><td>0–49</td><td>1%</td></tr><tr><td>50–64</td><td>5%</td></tr><tr><td>65–79</td><td>29%</td></tr><tr><td>80–100</td><td>65%</td></tr></tbody>\n          </table>\n          <p>By category, the median store earned 92% of the available points for access, 94% for readability and 75% for navigability. Getting in and reading pages is mostly solved. Finding the way around is where agents struggle.</p>\n          <h2>Finding 1: Most stores let AI assistants in</h2>\n          <p>The good news: robots.txt rules blocked AI assistants or AI search on only 3% of stores. Counting every way we test, 14% of stores blocked or challenged AI assistants somewhere:</p>\n          <ul>\n            <li><strong>9%</strong> blocked or challenged a request that identified itself as an AI assistant, such as ChatGPT-User or Claude-User, on the home page, while the same request as a normal browser got straight through.</li>\n            <li><strong>8%</strong> kept AI assistants or AI search out of product, pricing or cart pages, through bot protection or robots.txt.</li>\n            <li><strong>5%</strong> had a CAPTCHA on the home page, which agents can't solve.</li>\n          </ul>\n          <p>One caution on the bot protection figures: our requests used the AI assistants' names but came from our own servers, not from the operators' verified addresses. A site that checks where AI agents really come from would rightly turn some of these away, so these are upper bounds on how often real assistants are blocked. The robots.txt figure reads the rules each site publishes, so it isn't affected.</p>\n          <p>For the stores that do block assistants, the fix is usually a settings change, not a rebuild. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a> and <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>.</p>\n          <h2>Finding 2: Product details are where readability slips</h2>\n          <p>Most AI crawlers and assistant fetchers read the HTML a server sends and never run JavaScript. Here most stores do well: only 6% had less than 70% of their home page text in the HTML.</p>\n          <p>Product pages are a different story:</p>\n          <ul>\n            <li><strong>63%</strong> of stores give price, currency and availability as structured data in the product page's HTML, the way agents read most reliably.</li>\n            <li><strong>24%</strong> have no usable product structured data at all, so agents have to guess price and stock from the layout.</li>\n            <li><strong>6%</strong> have product data that's incomplete or only in older microdata, and 7% only add it with JavaScript.</li>\n            <li><strong>30%</strong> don't have prices in the product page's HTML at all.</li>\n          </ul>\n          <p>An agent that can't see a price can't compare you, and an agent that can't tell whether something is in stock will usually recommend a store where it can. See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data AI shopping agents can read</a>.</p>\n          <h2>Finding 3: Navigation is the weakest link</h2>\n          <p>Navigability was the lowest-scoring category, and the most common problems in the whole study are here:</p>\n          <ul>\n            <li><strong>48%</strong> of stores have buttons or links with no name an agent can read, most often icon-only cart, search and menu buttons.</li>\n            <li><strong>47%</strong> have elements that look clickable but aren't real buttons or links.</li>\n            <li><strong>39%</strong> show a pop-up or banner that covers much of the page on arrival. On 28% of stores, it had no clearly labelled button to close it.</li>\n            <li>Of 118 stores where we found the add-to-cart button, <strong>24%</strong> had one that isn't a real button, or is covered by something else, so an agent's click may not land.</li>\n          </ul>\n          <p>These are the same problems that trip up people using screen readers, and fixing them helps both. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>.</p>\n          <h2>Finding 4: Checkouts mostly let agents through</h2>\n          <p>AgentScore also loads each store's cart and checkout pages, without adding anything to the cart. Of 200 stores whose cart or checkout we could load, 8% blocked or challenged requests identifying as AI assistants there while browsers got through, and another 14% showed a CAPTCHA to every visitor. The same caution about verified addresses applies.</p>\n          <p>Most checkouts only show their first step once something is in the cart, so we could inspect too few of them to report on guest checkout or checkout forms. See <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">how to make checkout work for AI shopping agents</a> for what to look for in your own.</p>\n          <h2>Finding 5: More stores speak to agents than we expected</h2>\n          <p>Some stores now offer agents a direct way in. 53% publish an <a href=\"https://ghostagentlab.com/articles/llms-txt/\">llms.txt</a> file, 32% offer an MCP server or WebMCP tools, and 29% publish an agentic checkout profile. These standards are young, so we expected far lower numbers. They suggest agent-ready commerce is arriving faster than most store owners realize.</p>\n          <h2>The most common problems</h2>\n          <p>Across every check that covered at least 50 stores, these fell short most often (a fail or a warning):</p>\n          <table>\n            <thead><tr><th>AgentScore check</th><th>Stores falling short</th></tr></thead>\n            <tbody><tr><td>Buttons and links have names</td><td>48%</td></tr><tr><td>Clickable things are real buttons and links</td><td>47%</td></tr><tr><td>llms.txt guide for AI</td><td>47%</td></tr><tr><td>Pop-ups and banners can be dismissed by agents</td><td>39%</td></tr><tr><td>Product pages give price and stock in a form agents can read</td><td>37%</td></tr><tr><td>Form fields are labelled</td><td>34%</td></tr><tr><td>Prices are in the page HTML</td><td>30%</td></tr><tr><td>Page has main and navigation landmarks</td><td>28%</td></tr><tr><td>Clear page title, description, and headings</td><td>24%</td></tr><tr><td>AI agents aren't blocked or shown a CAPTCHA at the cart and checkout</td><td>22%</td></tr></tbody>\n          </table>\n          <h2>What the best stores do differently</h2>\n          <p>We compared the top quarter of stores by AgentScore with the bottom quarter. These checks showed the biggest gap in pass rates:</p>\n          <table>\n            <thead><tr><th>Check</th><th>Top quarter pass</th><th>Bottom quarter pass</th></tr></thead>\n            <tbody><tr><td>Product pages give price and stock in a form agents can read</td><td>100%</td><td>27%</td></tr><tr><td>Agents can check out through an agentic commerce protocol</td><td>72%</td><td>0%</td></tr><tr><td>llms.txt guide for AI</td><td>86%</td><td>16%</td></tr><tr><td>Agents can use your site through MCP or WebMCP</td><td>71%</td><td>2%</td></tr><tr><td>AI agents aren't blocked or shown a CAPTCHA at the cart and checkout</td><td>98%</td><td>41%</td></tr><tr><td>Prices are in the page HTML</td><td>94%</td><td>40%</td></tr></tbody>\n          </table>\n          <p>The leaders describe their products in a form agents can read, put prices in the HTML, publish an llms.txt file and offer agents a direct way in. Several of these are a one-time template or settings change, which makes them the cheapest points available to most stores.</p>\n          <h2>What to fix first</h2>\n          <div>\n            <ol>\n              <li><strong>Name your buttons, and make them real.</strong> Especially icon-only cart, search and menu buttons, and the add-to-cart button.</li>\n              <li><strong>Put product data and prices in the HTML.</strong> Name, price, currency and availability as structured data, rendered on the server.</li>\n              <li><strong>Get pop-ups out of the way</strong> on arrival, and give them a clearly labelled close button.</li>\n              <li><strong>Check your door.</strong> If you're among the stores that block AI assistants, let verified assistants reach your home, product and checkout pages.</li>\n            </ol>\n          </div>\n          <h2>See where your store stands</h2>\n          <p>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> on your own site. It's free, takes about two minutes, and shows exactly which of these problems apply to you, with a fix for each. To understand the bigger picture, start with <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">What is agent readiness?</a></p>\n          <p>Method notes: scores use AgentScore scoring version 2026.10-v5. Percentages use the stores where each check could run; checks that didn't apply to a store, or couldn't be tested, are left out of that check's total. Scans ran on October 9, 2026, and sites change, so a store's score today may differ. Checks that read what buttons and links say, such as \"Add to cart\" or \"Close\", understand 19 languages, so stores in every language count toward every figure.</p>",
      "date_published": "2026-10-09T13:23:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Research"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/what-an-ai-agent-sees/",
      "url": "https://ghostagentlab.com/blog/what-an-ai-agent-sees/",
      "title": "What an AI agent sees when it visits your store",
      "summary": "One product page, three views: the raw HTML an assistant reads, the rendered page a browser agent sees, and the accessibility tree it acts on.",
      "content_html": "<p>When your team reviews a product page, they look at it in a browser, logged in, with the cookie banner long since accepted. An AI agent arriving for a customer gets none of that. Depending on the kind of agent, it reads your raw HTML, looks at the page as it's drawn, or works through the page's accessibility tree. Here is one fictional product page seen all three ways, and what each view gets wrong.</p>\n          <h2>The page</h2>\n          <p>Northwind Outfitters sells running gear at <code>northwind.example</code>. Its product page for the Trail Runner 2 looks good to a person: a large photo, the name, a price of $129, a row of size swatches, an \"Add to cart\" button and, on a first visit, a cookie banner across the bottom half of the screen. Like many modern stores, the page is built in the browser: the server sends a shell, and JavaScript fetches the product and fills it in.</p>\n          <p>A shopper asks an AI assistant: \"Find me a trail running shoe under $150 in a size 10 that I can get this week.\" Here is what Northwind's page looks like to the agents that might come to answer.</p>\n          <h2>View 1: the raw HTML</h2>\n          <p>AI assistants that fetch a page because someone just asked a question, and the crawlers behind AI search, usually request the page and read the HTML that comes back. Many don't run JavaScript at all. This is what Northwind's server sends:</p>\n          <pre><code>&lt;html&gt;\n&lt;head&gt;\n  &lt;title&gt;Northwind Outfitters&lt;/title&gt;\n&lt;/head&gt;\n&lt;body&gt;\n  &lt;header&gt;&lt;a href=\"/\"&gt;&lt;img src=\"/logo.svg\"&gt;&lt;/a&gt;&lt;/header&gt;\n  &lt;div id=\"app\"&gt;Loading…&lt;/div&gt;\n  &lt;script src=\"/assets/app.4f9c.js\"&gt;&lt;/script&gt;\n&lt;/body&gt;\n&lt;/html&gt;</code></pre>\n          <p>That's the whole story for this kind of agent. No product name, because the title is the store's name. No price, no sizes, no stock and no structured data. The assistant can't tell the shopper the Trail Runner 2 costs $129 or comes in a size 10, so it recommends a shoe from a store whose page it could read.</p>\n          <p>Nothing on screen warns Northwind's team about this. The page looks perfect in every browser they own. It's the problem our guide to <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">JavaScript-only content</a> is about, and the reason <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">Product structured data</a> and <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">machine-readable prices</a> belong in the HTML the server sends.</p>\n          <h2>View 2: the rendered page</h2>\n          <p>Browser agents run a real browser. They load the page, let the scripts run, and decide what to do next, often by looking at a screenshot. This agent gets much more:</p>\n          <ul>\n            <li>The photo, the name \"Trail Runner 2\" and the price, $129.</li>\n            <li>The size swatches, though nothing on screen says which sizes are in stock: sold-out sizes are only a lighter shade of gray.</li>\n            <li>The cookie banner, covering the bottom of the screen, including the \"Add to cart\" button.</li>\n          </ul>\n          <p>The agent has to deal with the banner first. If its buttons are clearly labelled \"Accept\" and \"Reject\", that's one extra step. If the only way out is a small \"×\" drawn in an image, the agent may click the wrong thing or give up. Either way, it costs a step on every visit where the agent starts with a fresh browser, and many do. Our guide to <a href=\"https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/\">cookie banners and pop-ups</a> covers how to keep them out of the way.</p>\n          <h2>View 3: the accessibility tree</h2>\n          <p>Many browser agents don't act on pixels alone. They read the page's accessibility tree: the same outline of headings, buttons, links and fields that screen readers use, with a role and a name for each item. It's compact, and it tells the agent exactly what it can press. Here is a simplified version of Northwind's:</p>\n          <pre><code>banner\n  link (no name)\n  button (no name)\nmain\n  heading \"Trail Runner 2\" level 1\n  text \"$129.00\"\n  generic \"8\"\n  generic \"9\"\n  generic \"10\"\n  generic \"11\"\n  button \"Add to cart\"\ndialog \"We value your privacy\"\n  button \"Accept all\"\n  button (no name)</code></pre>\n          <p>Read it the way an agent would:</p>\n          <ul>\n            <li><strong>The logo link and the cart icon have no names.</strong> The agent can't tell which button is the cart. An <code>aria-label</code> or visible text fixes both.</li>\n            <li><strong>The sizes aren't controls.</strong> They're plain boxes with a click handler, so they show up as \"generic\" rather than buttons or options. The agent may not realize it can choose a size at all, and nothing tells it which sizes are sold out. Real buttons or radio inputs, with the stock state in their names, solve it. See <a href=\"https://ghostagentlab.com/articles/variant-pickers-ai-agents/\">variant pickers AI agents can use</a>.</li>\n            <li><strong>The banner's \"Reject\" button has no name.</strong> It's an icon with no label. An agent that wants to decline cookies on the shopper's behalf can't find the option.</li>\n          </ul>\n          <p>The good news is in there too: one clear <code>h1</code>, a price as text, and an \"Add to cart\" button that is a real button with a real name. Those are the things that let an agent finish the job.</p>\n          <h2>Same page, three stories</h2>\n          <table>\n            <thead><tr><th>What the agent needs</th><th>Raw HTML</th><th>Rendered page</th><th>Accessibility tree</th></tr></thead>\n            <tbody>\n              <tr><td>Product name</td><td>Missing</td><td>Yes</td><td>Yes</td></tr>\n              <tr><td>Price</td><td>Missing</td><td>Yes</td><td>Yes, as text</td></tr>\n              <tr><td>Sizes in stock</td><td>Missing</td><td>Only as shading</td><td>Sizes listed, stock missing, not selectable</td></tr>\n              <tr><td>A way past the banner</td><td>Not applicable</td><td>Depends on the buttons</td><td>\"Accept all\" only</td></tr>\n              <tr><td>Add to cart</td><td>Missing</td><td>Covered by the banner</td><td>Present and named</td></tr>\n            </tbody>\n          </table>\n          <p>No single view is \"the\" agent view. A shopper's question might be answered from the raw HTML, and the purchase made by a browser agent working from the accessibility tree. A page has to work in all three.</p>\n          <h2>How to see your own pages this way</h2>\n          <ol>\n            <li><strong>Raw HTML:</strong> open a product page, choose \"View page source\", and search for the price and the product name. If they're not there, agents that read the HTML don't have them.</li>\n            <li><strong>Rendered page:</strong> open the same page in a private window at desktop size. What covers it on arrival is what a browser agent meets first.</li>\n            <li><strong>Accessibility tree:</strong> in Chrome's developer tools, the Accessibility pane shows each element's role and name. Look for buttons and links with no name and for clickable things that show up as \"generic\".</li>\n          </ol>\n          <p><a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> does this comparison for you. It fetches your home page and key pages as a browser and as AI assistants, renders them in a real browser, and checks names, labels, banners and the add-to-cart button. Checks such as <strong>Content loads without JavaScript</strong>, <strong>Product pages give price and stock in a form agents can read</strong>, <strong>Buttons and links have names</strong> and <strong>Pop-ups and banners can be dismissed by agents</strong> map directly onto the problems above. <a href=\"https://ghostagentlab.com/blog/how-agentscore-renders-pages/\">How AgentScore renders your pages</a> explains the details.</p>\n          <div>\n            <p><strong>For Northwind, four changes cover most of it:</strong> render the product name, price and stock in the HTML with Product structured data; give the logo, cart icon and banner buttons readable names; make the size swatches real controls that say when a size is sold out; and keep the banner off the \"Add to cart\" button. None of them changes how the page looks to a person. As we've argued before, <a href=\"https://ghostagentlab.com/blog/accessibility-is-agent-readiness/\">accessibility work is agent readiness work</a>.</p>\n          </div>",
      "date_published": "2026-10-09T13:22:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Perspective"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/holiday-season-agent-readiness/",
      "url": "https://ghostagentlab.com/blog/holiday-season-agent-readiness/",
      "title": "Getting agent-ready before the holiday rush",
      "summary": "Bot rules, sale prices, stock, promo pop-ups and code freezes: what to check before peak season so AI agents can still find, price and buy.",
      "content_html": "<p>Peak season is when a broken journey costs the most, and it's also when the riskiest changes go live: tighter bot rules, sale prices, promo pop-ups, then a code freeze that makes anything you missed hard to fix. Here's what to check before the freeze, so AI agents shopping for your customers can still find you, read your prices and buy.</p>\n          <h2>Why peak season is different</h2>\n          <p>Every change made for peak is reasonable on its own. Security tightens bot protection to stop scalpers and card testing. Marketing adds a countdown banner and an email pop-up. Merchandising loads sale prices through a promotions app. Operations shortens cache times, or lengthens them to cope with load. Then engineering freezes the code.</p>\n          <p>Each of those changes is made by a different team, often in a hurry, and none of them is tested with an AI agent in mind. An assistant asked to \"find a gift under $50 that arrives before the 24th\" needs to get past your bot rules, read the sale price, see that the item is in stock, close the pop-up and reach the cart. Any one of those can fail quietly, and you won't see it in your analytics.</p>\n          <h2>A timeline for the weeks before the freeze</h2>\n          <table>\n            <thead><tr><th>When</th><th>What to do</th><th>Who</th></tr></thead>\n            <tbody>\n              <tr><td>Six or more weeks out</td><td>Run AgentScore on your home page and fix anything critical. Agree which journeys matter most at peak.</td><td>E-commerce lead, developers</td></tr>\n              <tr><td>Four weeks out</td><td>Review planned bot protection changes with your CDN or security team. Check sale price and stock data on a test promotion.</td><td>Security, merchandising</td></tr>\n              <tr><td>Two weeks out</td><td>Put promo pop-ups and banners live on staging or behind a flag, and check an agent can still reach the cart.</td><td>Marketing, developers</td></tr>\n              <tr><td>Before the freeze</td><td>Rescan, set up scheduled journey tests with alerts, and agree what counts as an exception to the freeze.</td><td>E-commerce lead</td></tr>\n            </tbody>\n          </table>\n          <h2>1. Bot protection tightened for peak</h2>\n          <p>This is the change most likely to turn agents away. Common peak settings include stricter challenge thresholds, \"under attack\" modes that challenge every visitor, blocks on cloud and data center networks, and country blocks. Browser agents and assistants that fetch pages for people often run from cloud networks, so they get caught by rules written for scrapers.</p>\n          <ul>\n            <li>Ask for a list of every bot rule planned for peak, and who can change them during the freeze. CDN rules often sit outside the code freeze and change mid-season.</li>\n            <li>Make sure the AI assistants you want are allowed by verified identity, not just by user agent. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</li>\n            <li>Check cart and checkout separately from the home page. Many stores add extra protection there for peak.</li>\n            <li>Prefer rate limits on the expensive actions, like payment attempts, over blanket challenges. See <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a> and <a href=\"https://ghostagentlab.com/articles/geo-blocking-ai-agents/\">geo-blocking and AI agents</a>.</li>\n          </ul>\n          <p>AgentScore checks whether your bot protection treats AI assistants differently from a browser, and whether the cart and checkout show them a block or a CAPTCHA. Rescan after every bot rule change, not just once in October. For the settings themselves, see <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a> and <a href=\"https://ghostagentlab.com/articles/cdn-settings-ai-agents/\">setting up your CDN for AI agents</a>.</p>\n          <h2>2. Sale prices that agents can read</h2>\n          <p>Promotion tools often change the price in the browser with JavaScript after the page loads. A shopper sees the sale price. An agent that reads the HTML sees the full price, or no price at all, and may quote the wrong figure to its user or skip you.</p>\n          <p>Before peak, run a test promotion and check three places: the visible price in the page source (use View Source, not the inspector), the price in your product structured data, and the price in the cart. They should all agree. If a sale ends on a known date, say so:</p>\n          <pre><code>\"offers\": {\n  \"@type\": \"Offer\",\n  \"price\": \"39.00\",\n  \"priceCurrency\": \"USD\",\n  \"priceValidUntil\": \"2026-12-01\",\n  \"availability\": \"https://schema.org/InStock\"\n}</code></pre>\n          <p>Show the original and the sale price in text, such as \"Was $59, now $39\", rather than relying on a strikethrough alone. See <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">prices AI agents can read</a>.</p>\n          <h2>3. Stock that stays accurate</h2>\n          <p>At peak, stock changes by the hour. If your structured data says <code>InStock</code> after an item sells out, an assistant will send people to a product they can't buy. If it says <code>OutOfStock</code> for an item you restocked this morning, you lose the recommendation.</p>\n          <ul>\n            <li>Make sure <code>availability</code> in your product data comes from the same source as the \"Add to cart\" button, not a nightly export.</li>\n            <li>Use <code>PreOrder</code> or <code>BackOrder</code> where they apply, and state delivery estimates in text.</li>\n            <li>If you lengthen CDN cache times for load, check how stale product pages can get, and purge on stock changes for best sellers.</li>\n          </ul>\n          <h2>4. Promo pop-ups and banners</h2>\n          <p>Email capture, spin-to-win wheels, countdown banners and free-shipping bars all multiply in November. An agent can only get past one if it has a real, labelled close button and doesn't cover the content behind it. A pop-up that sits over the \"Add to cart\" button means the agent's click lands on the pop-up instead.</p>\n          <p>AgentScore checks that pop-ups and banners can be dismissed, and whether the add-to-cart button on a product page is covered by something else. Check again once peak promotions are live, because they're usually added after the last scan. See <a href=\"https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/\">cookie banners and pop-ups</a>.</p>\n          <h2>5. Promo codes, carts and checkout</h2>\n          <p>Gift shopping by agent depends on the last few steps working. Make sure the promo code field is labelled, that a rejected code produces a clear error in text, and that the cart total updates in the page. See <a href=\"https://ghostagentlab.com/articles/cart-ai-agents/\">carts AI agents can manage</a> and <a href=\"https://ghostagentlab.com/articles/error-messages-ai-agents/\">error messages AI agents can understand</a>.</p>\n          <p>Resist adding a CAPTCHA to checkout for peak. It stops agents at the very last step, after they've done all the work. Card testing is better handled by your payment provider's fraud tools and by rate limits on payment attempts. Keep guest checkout on. See <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">checkout for AI shopping agents</a> and <a href=\"https://ghostagentlab.com/articles/guest-checkout-ai-agents/\">guest checkout</a>.</p>\n          <h2>6. Holiday policies in plain text</h2>\n          <p>Assistants are often asked about delivery cut-off dates and extended returns. Put them on your shipping and returns pages as text in the HTML, with dates, not only in a banner image or a pop-up. See <a href=\"https://ghostagentlab.com/articles/policy-pages-ai-assistants/\">policy pages AI assistants can read</a>.</p>\n          <h2>Planning for the freeze itself</h2>\n          <p>A code freeze protects you from risky changes, but it also means a problem found in December may wait until January. Three things help:</p>\n          <ol>\n            <li><strong>Agree that blocking AI agents counts as an incident.</strong> If an assistant can't reach your product pages or checkout, that should qualify for a freeze exception, the same as a broken payment form.</li>\n            <li><strong>Test the money journeys on a schedule.</strong> <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> run real journeys, such as search, choose a size and add to cart, stop at the payment form, and alert you when a journey that used to pass starts failing, with a step-by-step replay.</li>\n            <li><strong>Watch your logs, not just analytics.</strong> Most agents don't run your analytics tag. A rise in blocked or failed agent requests shows up in server and CDN logs first. See <a href=\"https://ghostagentlab.com/articles/agent-errors-in-logs/\">finding the errors AI agents hit</a>.</li>\n          </ol>\n          <div>\n            <p><strong>Before the freeze, check:</strong></p>\n            <ul>\n              <li>Every peak bot rule has been reviewed, and verified AI assistants still get through to product pages, cart and checkout.</li>\n              <li>Sale prices appear in the page source and in structured data, and match the cart.</li>\n              <li>Stock status in product data updates when items sell out.</li>\n              <li>Every promo pop-up has a labelled close button and doesn't cover \"Add to cart\".</li>\n              <li>No new CAPTCHA at checkout. Guest checkout is on.</li>\n              <li>Delivery cut-offs and holiday returns are written out on your policy pages.</li>\n              <li>Scheduled journey tests and alerts go to someone who can act during the freeze.</li>\n            </ul>\n          </div>\n          <p>Start with a fresh <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> scan, then work through the list. If you have more time, our <a href=\"https://ghostagentlab.com/blog/90-day-agent-readiness-plan/\">90-day agent readiness plan</a> covers the rest.</p>",
      "date_published": "2026-10-09T13:21:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Guide"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/agent-readiness-for-saas/",
      "url": "https://ghostagentlab.com/blog/agent-readiness-for-saas/",
      "title": "Agent readiness for software companies: pricing, sign-up and docs",
      "summary": "For software companies, the pricing page is the product page and sign-up is the checkout. How to get pricing, trials, docs, llms.txt and MCP right.",
      "content_html": "<p>Most agent readiness advice is written for online stores. But software buyers send AI agents to do research too: compare plans, check whether a tool integrates with what they already use, find the limits of a free tier, and start a trial. For a software company, the pricing page is the product page, the sign-up form is the checkout, and the docs are where the real evaluation happens.</p>\n          <h2>How AgentScore sees a software site</h2>\n          <p>AgentScore looks for the pages an agent acting for a customer would need, the way an agent would: links in your home page's HTML, then your sitemap, then common addresses such as <code>/pricing</code>. If it finds product pages or a cart, it treats the site as a store. If it finds a pricing page and no products or cart, it treats the site as software (\"saas\" in the scan's data). That changes what some checks look for.</p>\n          <table>\n            <thead><tr><th>Check</th><th>On a software site, it looks at</th></tr></thead>\n            <tbody>\n              <tr><td>AI agents can open your product, pricing and cart pages</td><td>Whether AI assistants can load your pricing page, not just a browser</td></tr>\n              <tr><td>Prices are in the page HTML</td><td>Whether plan prices are in the HTML your server sends, before JavaScript runs</td></tr>\n              <tr><td>Key pages are linked from the home page</td><td>Whether pricing is linked from the home page with an ordinary link</td></tr>\n              <tr><td>Agents can find and press your add-to-cart or sign-up button</td><td>Whether the pricing page has a real, visible button such as \"Start free trial\" or \"Buy Pro\"</td></tr>\n              <tr><td>Cart, checkout and agentic commerce checks</td><td>Nothing. They're marked \"not tested\", so they don't count against you</td></tr>\n            </tbody>\n          </table>\n          <p>Everything else, from robots.txt and bot protection to structured data, named buttons and labelled form fields, applies the same way as for any site. See <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">what is agent readiness?</a> for the full model.</p>\n          <h2>Pricing pages agents can read</h2>\n          <p>Pricing pages are often the most designed page on a software site, and that's where agents struggle. Common problems:</p>\n          <ul>\n            <li><strong>Prices loaded by JavaScript.</strong> A monthly/annual toggle that fetches prices when clicked leaves the HTML with no prices at all. Put both sets of prices in the HTML and let the toggle show and hide them.</li>\n            <li><strong>Features shown only as icons.</strong> A comparison table of tick and cross images says nothing to an agent unless each icon has a text label, such as \"Included\" or \"Not included\". See <a href=\"https://ghostagentlab.com/articles/image-alt-text-ai-agents/\">image alt text</a>.</li>\n            <li><strong>Units left implicit.</strong> Say \"per user per month, billed annually\" in text, not in a tooltip. Agents compare plans across vendors, and unclear units make you look more expensive or less clear than a competitor.</li>\n            <li><strong>\"Contact sales\" with no explanation.</strong> If a plan is priced on request, say so in plain text and say what affects the price. An agent can then tell its user how to get a quote instead of guessing.</li>\n          </ul>\n          <p>A plain HTML table is still the clearest way to compare plans:</p>\n          <pre><code>&lt;table&gt;\n  &lt;caption&gt;Northwind plans, billed annually&lt;/caption&gt;\n  &lt;tr&gt;&lt;th&gt;&lt;/th&gt;&lt;th&gt;Starter&lt;/th&gt;&lt;th&gt;Pro&lt;/th&gt;&lt;/tr&gt;\n  &lt;tr&gt;&lt;th&gt;Price&lt;/th&gt;&lt;td&gt;$12 per user per month&lt;/td&gt;&lt;td&gt;$29 per user per month&lt;/td&gt;&lt;/tr&gt;\n  &lt;tr&gt;&lt;th&gt;Single sign-on&lt;/th&gt;&lt;td&gt;Not included&lt;/td&gt;&lt;td&gt;Included&lt;/td&gt;&lt;/tr&gt;\n&lt;/table&gt;</code></pre>\n          <p>For more, see <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">prices AI agents can read</a>.</p>\n          <h2>Sign-up and trials</h2>\n          <p>An agent that has chosen your product for someone will try to start the trial. It needs a visible, real button with a clear label, then a form it can fill in: every field labelled, errors explained in text, and no surprises. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>.</p>\n          <p>Some steps should stay with the person, and that's fine. Email verification, single sign-on, payment details for a card-required trial and accepting terms are all reasonable places for an agent to hand back to its user. The aim is for the agent to get as far as it should, then stop cleanly with a clear message, rather than fail at a CAPTCHA on the first screen. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-sign-up-booking/\">sign-up, booking and lead forms AI agents can finish</a>.</p>\n          <p><a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> can test this journey for real: from the home page to pricing to the sign-up form, using test data you supply, with a step-by-step replay showing where an agent got stuck.</p>\n          <h2>Docs are your product page for agents</h2>\n          <p>For many software products, the docs decide the sale. Coding assistants and research agents read them to answer \"does it support X?\" and \"how hard is it to set up?\". Treat docs with the same care as marketing pages:</p>\n          <ul>\n            <li>Serve docs as HTML that works without JavaScript. Some docs sites are single-page apps that send an empty shell. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">JavaScript-only content</a>.</li>\n            <li>Keep public docs public. If setup guides sit behind a login, agents can't use them to recommend you. See <a href=\"https://ghostagentlab.com/articles/login-walls-ai-agents/\">login walls and gated content</a>.</li>\n            <li>Give each version a stable address and point old copies at the current one, so agents don't quote outdated instructions. See <a href=\"https://ghostagentlab.com/articles/canonical-urls-ai-agents/\">canonical URLs</a>.</li>\n            <li>Publish limits, supported integrations and security information as text, not only in PDFs or sales decks.</li>\n          </ul>\n          <h2>llms.txt for software companies</h2>\n          <p><a href=\"https://llmstxt.org/\">llms.txt</a> is a proposed convention, not a standard, and AI providers don't all say whether they read it. It's cheap to publish, though, and it suits software companies well, because your most useful pages are easy to list. Some docs platforms can generate one for you.</p>\n          <pre><code># Northwind\n&gt; Northwind is scheduling software for field service teams.\n## Product\n- [Pricing](https://northwind.example/pricing): plans, limits and what each includes\n- [Start a free trial](https://northwind.example/signup): 14 days, no card required\n## Docs\n- [Getting started](https://northwind.example/docs/start)\n- [API reference](https://northwind.example/docs/api)\n- [Integrations](https://northwind.example/integrations)\n## Trust\n- [Security](https://northwind.example/security)\n- [Status](https://status.northwind.example/)</code></pre>\n          <p>AgentScore checks that an llms.txt file exists. It can't tell whether yours is up to date, so add it to your release checklist. See <a href=\"https://ghostagentlab.com/articles/llms-txt/\">how to write an llms.txt file</a>.</p>\n          <h2>MCP: your product, not just your website</h2>\n          <p>For a software company, the <a href=\"https://modelcontextprotocol.io/\">Model Context Protocol</a> is about more than your website. An MCP server can let an assistant use your product itself on a customer's behalf: look up records, create tasks, run reports. That's product work, with sign-in and permissions to design carefully, and many software companies are already building one.</p>\n          <p>There's no single agreed way to advertise an MCP server on a website yet. AgentScore's \"Agents can use your site through MCP or WebMCP\" check looks for a server card at <code>/.well-known/mcp.json</code>, an MCP link in the page, a mention in llms.txt, or WebMCP tools in the page. If you have a server, mentioning it in llms.txt and your docs is the simplest start. See <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP and WebMCP for websites</a>.</p>\n          <h2>Where to start</h2>\n          <ol>\n            <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> on your home page. If \"Key pages are linked from the home page\" doesn't pass, agents may not be finding your pricing page.</li>\n            <li>View the source of your pricing page and check every price and plan feature is there as text.</li>\n            <li>Make \"Start free trial\" a real, clearly labelled button, and check the sign-up form's fields are labelled.</li>\n            <li>Publish an llms.txt that links to pricing, sign-up, docs and your trust pages.</li>\n            <li>Test the journey from home page to sign-up with a <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a>.</li>\n          </ol>",
      "date_published": "2026-10-09T13:20:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Perspective"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/explain-agent-readiness-to-your-board/",
      "url": "https://ghostagentlab.com/blog/explain-agent-readiness-to-your-board/",
      "title": "How to explain agent readiness to your board",
      "summary": "A one-page way to explain agent readiness to your board: what's changing, the risks, what to measure and what to ask for, without hype or invented numbers.",
      "content_html": "<p>Your board doesn't need to know what robots.txt is. It needs to know what's changing, what that risks for the business, what you're doing about it, and how it will know the work is paying off. Here's a way to say all of that on one page, without hype and without numbers you can't stand behind.</p>\n          <h2>The one-page version</h2>\n          <p>Start with this structure. Each heading is one or two sentences. Fill in the bracketed parts with your own findings.</p>\n          <div>\n            <p><strong>What's changing.</strong> People are starting to ask AI assistants to research, compare and buy on their behalf. Those assistants visit our website the way a customer would, but they read it differently.</p>\n            <p><strong>Why it matters to us.</strong> If an assistant can't get in, can't read our prices or can't complete a purchase, it recommends someone else, and we never see the lost customer in our reports.</p>\n            <p><strong>Where we stand.</strong> [Our AgentScore, the main issues it found, and whether AI agents can complete our most important journey today.]</p>\n            <p><strong>What we're doing.</strong> [Three to five fixes, who owns each, and when they'll be done.]</p>\n            <p><strong>What we need.</strong> [Time, tools or decisions, each with a review date.]</p>\n          </div>\n          <p>That's it. Everything below is about filling it in well.</p>\n          <h2>Explaining the shift without hype</h2>\n          <p>Boards have heard a lot about AI. The quickest way to lose the room is a slide of large market forecasts. You don't need them. The argument for agent readiness is modest and easy to defend:</p>\n          <ul>\n            <li>AI agents already visit websites for people. You can show this from your own server or CDN logs.</li>\n            <li>Nobody knows how fast this will grow. That's a reason to be ready cheaply now, not to wait for certainty.</li>\n            <li>Much of the work also improves accessibility, search and site quality, so it isn't a bet on one trend. See <a href=\"https://ghostagentlab.com/blog/accessibility-is-agent-readiness/\">accessibility work is agent readiness work</a> and <a href=\"https://ghostagentlab.com/blog/seo-to-agent-readiness/\">SEO got you found. Agent readiness gets you chosen</a>.</li>\n            <li>The failure is silent. A blocked agent doesn't complain. It goes elsewhere. That makes it the kind of risk a board should want someone to own.</li>\n          </ul>\n          <p>If you quote any outside figure, cite the source on the slide. If you can't, describe the trend in words instead.</p>\n          <h2>The risks, in board language</h2>\n          <table>\n            <thead><tr><th>Risk</th><th>What it looks like</th><th>Usually owned by</th></tr></thead>\n            <tbody>\n              <tr><td>Lost revenue</td><td>AI assistants are blocked by bot protection, can't see prices, or fail at checkout, so they recommend or buy from a competitor.</td><td>E-commerce or digital lead</td></tr>\n              <tr><td>Being misquoted</td><td>Assistants give customers the wrong price, stock or return policy because our pages are hard to read or out of date.</td><td>Marketing and content</td></tr>\n              <tr><td>Impostors</td><td>Scrapers pretend to be well-known AI agents to get past our defenses. Opening the door carelessly makes this worse.</td><td>Security</td></tr>\n              <tr><td>Content and legal</td><td>Decisions about AI training on our content, our terms of service, and responsibility for purchases an agent makes for a customer.</td><td>Legal, with marketing</td></tr>\n              <tr><td>Blind spots</td><td>Our analytics can't see most AI agents, so decisions are made on incomplete data.</td><td>Analytics or digital lead</td></tr>\n            </tbody>\n          </table>\n          <p>The point of the table is that this is a cross-team issue with no natural home. Ask the board to agree one accountable owner. See <a href=\"https://ghostagentlab.com/blog/who-owns-agent-readiness/\">who owns agent readiness?</a></p>\n          <h2>What to measure</h2>\n          <p>Pick a few measures you can report every quarter, and show the trend rather than a single figure.</p>\n          <ul>\n            <li><strong>AgentScore, by category.</strong> Access, Readability and Navigability tell you whether agents can get in, understand your pages and find their way around. The free scan doesn't test Task completion, so don't present its score as proof that agents can buy.</li>\n            <li><strong>Journey pass rate.</strong> Whether AI agents can complete your most important journeys, such as checkout or sign-up, from scheduled <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> tests.</li>\n            <li><strong>Agent traffic.</strong> How many requests come from AI assistants and search agents, how many are verified, and how many are impostors, from your logs. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure agent traffic</a>.</li>\n            <li><strong>Visits and orders from AI assistants.</strong> People who click through from an assistant's answer. See <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">tracking visits and sales from AI assistants</a>.</li>\n            <li><strong>Time to fix.</strong> How long a critical finding, such as AI agents being blocked, stays open.</li>\n          </ul>\n          <p>Be careful with attribution. You can show that more agents get through and more journeys pass. You usually can't prove exactly how much revenue that produced, and the board will trust you more if you say so. For a fuller set, see <a href=\"https://ghostagentlab.com/articles/agent-readiness-kpis/\">agent readiness KPIs</a>.</p>\n          <h2>What to ask for</h2>\n          <p>Keep the ask small and staged, with a review point after each stage. A typical shape:</p>\n          <ol>\n            <li><strong>A baseline.</strong> A scan, a look at agent traffic in your logs, and one tested journey. This needs a few days of people's time, not a project.</li>\n            <li><strong>Fixes.</strong> Developer and content time to work through the findings. Much of it overlaps with work already on the accessibility and SEO backlog, so say where it does.</li>\n            <li><strong>Ongoing testing and monitoring.</strong> Tools that keep watching after the fixes, because a theme update or a new bot rule can undo them.</li>\n            <li><strong>Decisions.</strong> Time from legal and security to settle policy questions such as AI training crawlers and verified agents.</li>\n            <li><strong>Optional experiments.</strong> Newer interfaces such as <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP and WebMCP</a> or <a href=\"https://ghostagentlab.com/articles/agentic-commerce-protocols/\">agentic commerce protocols</a>, framed as small trials with a date to decide whether to continue.</li>\n          </ol>\n          <p>Our <a href=\"https://ghostagentlab.com/blog/90-day-agent-readiness-plan/\">90-day agent readiness plan</a> maps well onto the first three stages.</p>\n          <h2>Questions you'll probably get</h2>\n          <h3>Are we letting AI companies take our content?</h3>\n          <p>No. AI training crawlers and the assistants that fetch pages for a customer use different names, so you can block one and allow the other. That's a policy choice to make once, with legal. See <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">should you block AI crawlers?</a></p>\n          <h3>Does this open us up to bots?</h3>\n          <p>Not if it's done properly. The goal is to let in the agents you want, checked by verified identity, and keep blocking the rest. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</p>\n          <h3>What happens if we do nothing?</h3>\n          <p>Probably nothing you'll notice, which is the problem. Some customers' assistants will fail on the site and choose a competitor, and it won't show up in any report you currently read.</p>\n          <p>To fill in \"Where we stand\", run a free <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> scan. It takes under a minute and gives you a score out of 100 and a list of what to fix.</p>",
      "date_published": "2026-10-09T13:19:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Guide"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/agent-readiness-mistakes/",
      "url": "https://ghostagentlab.com/blog/agent-readiness-mistakes/",
      "title": "Seven agent readiness mistakes to avoid",
      "summary": "Seven common agent readiness mistakes, from blocking AI assistants with training crawlers to measuring agents with JavaScript analytics, and how to avoid each.",
      "content_html": "<p>Most agent readiness problems aren't decisions anyone made. They're side effects of sensible work: a security rule, a design choice, a tracking setup. Here are seven mistakes that are easy to make and easy to miss, why each one matters, and what to do instead.</p>\n          <h2>1. Blocking AI assistants along with training crawlers</h2>\n          <p>Many sites decide not to let AI companies train on their content, which is a fair choice. The mistake is blocking everything with \"AI\" in its name, including the assistants that fetch a page because a customer asked a question right now. Those are different agents with different names: <code>GPTBot</code> collects training data, while <code>ChatGPT-User</code> and <code>OAI-SearchBot</code> fetch pages for ChatGPT users and search. Anthropic draws the same line between <code>ClaudeBot</code> and <code>Claude-User</code>.</p>\n          <pre><code># Opt out of AI training\nUser-agent: GPTBot\nUser-agent: ClaudeBot\nUser-agent: Google-Extended\nDisallow: /\n# Everyone else, including AI assistants and AI search\nUser-agent: *\nAllow: /</code></pre>\n          <p><strong>Instead:</strong> decide on training and on assistants separately. See <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">should you block AI crawlers?</a> and <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>. The same split applies to bot protection rules, which often block by category rather than by name.</p>\n          <h2>2. Putting a CAPTCHA at checkout</h2>\n          <p>A CAPTCHA at checkout is often added to stop card testing, and it works on bots. It also stops every AI agent at the very last step, after it has searched, compared, chosen a size and filled the cart. That's the most expensive place to lose a customer.</p>\n          <p><strong>Instead:</strong> use your payment provider's fraud tools and rate limits on payment attempts, and keep challenges for traffic that is actually suspicious. AgentScore checks whether AI agents are blocked or shown a CAPTCHA at the cart and checkout. See <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">checkout for AI shopping agents</a> and <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a>.</p>\n          <h2>3. Prices that only appear with JavaScript</h2>\n          <p>Many sites build product and pricing pages in the browser. A person sees the price a moment after the page loads. Many AI agents read the HTML your server sends and never run the scripts, so they see a product with no price, or a placeholder. They may skip you, or tell their user to check the site.</p>\n          <p><strong>Instead:</strong> make sure prices are in the HTML, through server-side rendering or static generation, and in your product structured data. Check by viewing the page source, not the browser's inspector, which shows the page after scripts run. See <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">prices AI agents can read</a> and <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">JavaScript-only content</a>.</p>\n          <h2>4. Icons with no names</h2>\n          <p>A magnifying glass for search, a bag for the cart, a cross to close a pop-up. People know what they mean. Browser agents choose what to click by each control's name, the same way screen readers do, and an icon with no text or label has no name at all. To the agent, the cart button isn't there.</p>\n          <pre><code>&lt;button aria-label=\"Open cart\"&gt;\n  &lt;svg aria-hidden=\"true\"&gt;…&lt;/svg&gt;\n&lt;/button&gt;</code></pre>\n          <p><strong>Instead:</strong> give every icon-only button and link a label, and use real <code>&lt;button&gt;</code> and <code>&lt;a&gt;</code> elements rather than clickable boxes. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a> and <a href=\"https://ghostagentlab.com/blog/accessibility-is-agent-readiness/\">accessibility work is agent readiness work</a>.</p>\n          <h2>5. Publishing llms.txt and forgetting it</h2>\n          <p>llms.txt is a proposed convention for giving AI a short guide to your site. It's quick to write, which is also why it goes stale: a file written at launch still links to last year's collections, a retired plan, or a returns page that moved. An out-of-date guide can be worse than none, because it points agents at the wrong answer.</p>\n          <p><strong>Instead:</strong> give llms.txt an owner and add it to the checklist for site changes, alongside the sitemap. Note that AgentScore checks that the file exists, not whether what it says is still true, so that part is up to you. See <a href=\"https://ghostagentlab.com/articles/llms-txt/\">how to write an llms.txt file</a>.</p>\n          <h2>6. Trusting user agents</h2>\n          <p>A user agent is a name a visitor gives itself, and anyone can use any name. If you allow a list of AI agents by user agent alone, scrapers borrow those names to get in. If you block by user agent, you may turn away a real assistant while impostors simply change their name.</p>\n          <p><strong>Instead:</strong> verify identity. The large operators publish their IP ranges, and many crawlers can be confirmed with a reverse DNS lookup. Newer agents can sign their requests with Web Bot Auth, which your CDN may already check for you. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a> and <a href=\"https://ghostagentlab.com/blog/signed-requests/\">why we sign every request our agents send</a>. Ghost Agent Labs' Verification page checks every request that claims to be a known agent and shows the impostors.</p>\n          <h2>7. Measuring AI agents with JavaScript analytics</h2>\n          <p>Analytics tools such as Google Analytics count visitors by running a script in the browser. Most AI agents don't run it, so they never appear. A report that shows no AI agents isn't evidence that none came. It may mean they came and the tag didn't see them, or that they were blocked before the page loaded.</p>\n          <p><strong>Instead:</strong> measure agents from your server or CDN logs, which record every request whether or not scripts run. Keep analytics for the people who click through from an assistant's answer, which it can see. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">why your analytics can't see AI agents</a>, <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">AI referral traffic</a> and <a href=\"https://ghostagentlab.com/blog/agent-traffic-is-not-bot-traffic/\">agent traffic isn't bot traffic</a>.</p>\n          <h2>What these have in common</h2>\n          <p>None of these mistakes shows up from inside the business. The site looks fine in a browser, the dashboards look normal, and the security team sees fewer bad bots. The cost lands on customers using an assistant, who quietly go elsewhere.</p>\n          <table>\n            <thead><tr><th>Mistake</th><th>Quick check</th><th>Usually fixed by</th></tr></thead>\n            <tbody>\n              <tr><td>Assistants blocked with training crawlers</td><td>Read /robots.txt and your bot rules for AI categories</td><td>SEO, bot protection or CDN admin</td></tr>\n              <tr><td>CAPTCHA at checkout</td><td>Ask your security team what protects the checkout</td><td>Bot protection or CDN admin</td></tr>\n              <tr><td>JavaScript-only prices</td><td>View source on a product page and search for the price</td><td>Developer</td></tr>\n              <tr><td>Icons with no names</td><td>Tab through the header and listen with a screen reader</td><td>Developer</td></tr>\n              <tr><td>Stale llms.txt</td><td>Open /llms.txt and click every link</td><td>SEO or content team</td></tr>\n              <tr><td>Trusting user agents</td><td>Ask how \"allowed\" AI agents are identified</td><td>Bot protection or CDN admin</td></tr>\n              <tr><td>JavaScript analytics only</td><td>Ask where agent traffic numbers come from</td><td>Analytics or digital lead</td></tr>\n            </tbody>\n          </table>\n          <p>A free <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> scan catches several of these in under a minute. For the journeys themselves, <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> show whether an agent can actually finish them.</p>",
      "date_published": "2026-10-09T13:18:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Guide"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/how-agentscore-renders-pages/",
      "url": "https://ghostagentlab.com/blog/how-agentscore-renders-pages/",
      "title": "How AgentScore renders your pages in a real browser",
      "summary": "Which pages AgentScore opens in a headless browser, with what settings and safeguards, how it looks for WebMCP tools, and what it never does.",
      "content_html": "<p>Some AI agents read your HTML exactly as your server sends it. Others drive a real browser, run your scripts and look at the finished page. To score your site for both, AgentScore looks at your key pages twice: once as served, and once rendered in a headless browser. Here is exactly what that browser does, how we keep it safe, and what it deliberately never does.</p>\n          <h2>Why render at all</h2>\n          <p>Several checks only make sense as a comparison between the page before and after JavaScript. <strong>Content loads without JavaScript</strong> measures how much of the rendered page's text was already in the HTML: 70% or more passes, 30% to 70% is a warning, and less than that fails. <strong>Structured data describes your business and products</strong>, <strong>Product pages give price and stock in a form agents can read</strong> and <strong>Prices are in the page HTML</strong> all look for the same thing: data that only appears once scripts have run, which agents that read the HTML never see. Our guide to <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">JavaScript-only content</a> covers why that matters.</p>\n          <p>The Navigability checks need a rendered page too: whether a button is visible or a banner covers the screen depends on layout, which only exists once a browser has drawn the page.</p>\n          <h2>Which pages we render</h2>\n          <p>A scan renders at most four pages, each in its own browser session:</p>\n          <table>\n            <thead><tr><th>Page</th><th>When it's rendered</th><th>What we look at</th></tr></thead>\n            <tbody>\n              <tr><td>Home page</td><td>Every scan</td><td>The Navigability checks, image alt text, text and structured data after JavaScript, WebMCP tools</td></tr>\n              <tr><td>Product page</td><td>When we find one and it loads for a browser</td><td>The add-to-cart button, prices and product data after JavaScript, WebMCP tools</td></tr>\n              <tr><td>Pricing page</td><td>When we find one and it loads for a browser</td><td>The sign-up or buy button, prices after JavaScript, WebMCP tools</td></tr>\n              <tr><td>Checkout</td><td>Only when the HTML we were sent has no form fields and the checkout didn't send us back to the cart</td><td>The form fields a checkout builds with JavaScript</td></tr>\n            </tbody>\n          </table>\n          <p>The cart is fetched but not rendered. <a href=\"https://ghostagentlab.com/blog/how-agentscore-works/\">How AgentScore works</a> explains how key pages are found and lists every page a scan requests.</p>\n          <h2>The browser and its settings</h2>\n          <ul>\n            <li><strong>Chromium, headless.</strong> The engine behind Chrome, driven by Playwright.</li>\n            <li><strong>A fresh, isolated session for every page.</strong> No cookies or storage carry over, so every page sees a first-time visitor.</li>\n            <li><strong>A desktop screen, 1280 by 900 pixels,</strong> and a desktop Chrome user agent. These requests aren't signed, so your site treats them as it would any desktop visitor.</li>\n            <li><strong>No downloads and no service workers.</strong></li>\n            <li><strong>Up to 20 seconds to load.</strong> We wait for the HTML to be parsed, then up to 5 more seconds for network activity to settle, so content built in the browser and consent banners have time to appear. Pages with chat widgets or analytics beacons rarely go quiet, so after 5 seconds we inspect whatever is there.</li>\n          </ul>\n          <p>If a page fails to load or render, the checks that depend on it are marked \"not tested\" with a note explaining why. A render failure never counts as a fail.</p>\n          <h2>What we inspect once it's drawn</h2>\n          <p>On the home page, a script reads the finished page and records:</p>\n          <ul>\n            <li>Every visible button and link, and whether it has a name an agent can read: its text, an <code>aria-label</code>, a referenced label, a title, or the alt text of an image inside it.</li>\n            <li>Every visible form field, and whether a label is connected to it. Fields with only placeholder text are counted separately, as a warning.</li>\n            <li>The largest fixed or sticky element covering at least 15% of the screen, leaving out ordinary header bars along the top, and whether it has a clearly labelled button such as \"Accept\", \"Close\" or \"No thanks\".</li>\n            <li>Things that look clickable but aren't buttons or links, the <code>&lt;main&gt;</code> and <code>&lt;nav&gt;</code> landmarks, and images with no <code>alt</code> attribute.</li>\n          </ul>\n          <p>On product and pricing pages, we look for the main action by its text, such as \"Add to cart\", \"Buy now\", \"Start free trial\" or \"Sign up\". We scroll it into view, check whether it's a real button or link, and ask the browser what element sits at its center. If that's a cookie banner rather than the button, an agent's click would land on the banner, and <strong>Agents can find and press your add-to-cart or sign-up button</strong> fails.</p>\n          <h2>The WebMCP probe</h2>\n          <p>WebMCP is a proposal for a browser feature, <code>navigator.modelContext</code>, that lets a page offer tools directly to an AI agent in the browser, such as \"search products\" or \"add to cart\". It isn't generally available in browsers yet, so sites that support it check for it first, roughly like this:</p>\n          <pre><code>if (\"modelContext\" in navigator) {\n  navigator.modelContext.registerTool({\n    name: \"search_products\",\n    description: \"Search the Northwind catalog by keyword\",\n    inputSchema: { type: \"object\", properties: { query: { type: \"string\" } } },\n    execute: async ({ query }) =&gt; searchCatalog(query),\n  });\n}</code></pre>\n          <p>Before your scripts run, we add a stand-in for <code>navigator.modelContext</code> that writes down the name of every tool a page registers. A page that checks for WebMCP then registers its tools just as it would in a browser that supports it. We also look for forms marked up with a <code>toolname</code> attribute. We run the probe on the home, product and pricing pages and record names only: our stand-in never calls a tool. The result feeds <strong>Agents can use your site through MCP or WebMCP</strong>, where a missing setup is a warning, never a fail. Our guide to <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP and WebMCP for websites</a> explains both.</p>\n          <h2>Keeping the browser safe</h2>\n          <p>A service that opens any URL it's given is a target: someone could submit an address that points at an internal network or a cloud metadata service. So every connection the scanner makes is checked:</p>\n          <ol>\n            <li><strong>Only http and https.</strong> Other schemes are refused.</li>\n            <li><strong>Public addresses only.</strong> Every address a host name resolves to must be publicly routable. Private, loopback, link-local, shared, reserved, documentation and multicast ranges are refused, for IPv4 and IPv6, including IPv4 addresses hidden inside IPv6 ones. Cloud metadata addresses are always refused.</li>\n            <li><strong>Checked twice.</strong> Each request the page makes is checked before it goes out. Then all browser traffic, WebSockets included, passes through a small proxy that looks up the host name itself and connects only to addresses that pass. A host can't answer with a public address for the check and a private one for the connection.</li>\n            <li><strong>Limits on plain fetches.</strong> Requests made without the browser are checked again at every redirect, follow at most five redirects, time out after 15 seconds and stop reading at a size limit, 5 MB for an ordinary page.</li>\n          </ol>\n          <h2>What the browser never does</h2>\n          <ul>\n            <li>It never clicks, types, signs in, adds to cart or submits a form. The only interaction is scrolling the main button into view to see what's on top of it.</li>\n            <li>It doesn't dismiss your cookie banner. We measure the page as a first-time visitor meets it, because that's how most agents arrive.</li>\n            <li>It doesn't pretend to be an AI agent. Comparisons between agents use plain requests, as described in <a href=\"https://ghostagentlab.com/blog/scanner-wears-other-names/\">Why our scanner sometimes wears other agents' names</a>.</li>\n            <li>It doesn't test a mobile layout, and it can't see a checkout that only opens with items in the cart. Checks that depend on one say \"not tested\" rather than guess.</li>\n          </ul>\n          <p>Finishing a real task is a different job. That's what <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> do in the app: real AI agents run your journeys, stop at the payment form, and give you a step-by-step replay. The free scan marks Task completion as not tested and scores the other three categories.</p>\n          <div>\n            <p><strong>To reproduce a finding:</strong> open the page in a private window at desktop size and compare \"View source\" with what's on screen. If the price, product details or button is on screen but not in the source, agents that read the HTML don't get it. If a banner sits over the button on arrival, a browser agent meets it first.</p>\n          </div>",
      "date_published": "2026-10-09T13:17:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Engineering"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/reading-your-agentscore-report/",
      "url": "https://ghostagentlab.com/blog/reading-your-agentscore-report/",
      "title": "Reading your AgentScore report: what to fix first",
      "summary": "What each part of an AgentScore report means, from the verdict and statuses to severity and score impact, and how to turn it into a fix list.",
      "content_html": "<p>An AgentScore report gives you a number, a sentence and a list of checks. The number tells you where you stand. The list tells you what to do about it, as long as you read it in the right order. This guide walks through each part of the report, in the free scan and in the Ghost Agent Labs app, and shows how to turn it into a short list of fixes with an owner for each.</p>\n          <h2>Start at the top: score, band and verdict</h2>\n          <p>The score is out of 100. It falls into one of four bands, and the verdict under it says what the band means:</p>\n          <table>\n            <thead><tr><th>Score</th><th>Band</th><th>Verdict</th></tr></thead>\n            <tbody>\n              <tr><td>80 and up</td><td>Good</td><td>AI agents can use this site well.</td></tr>\n              <tr><td>60 to 79</td><td>Fair</td><td>AI agents can mostly use this site, but some things get in their way.</td></tr>\n              <tr><td>40 to 59</td><td>Weak</td><td>AI agents will struggle on this site.</td></tr>\n              <tr><td>Below 40</td><td>Poor</td><td>Most AI agents will fail on this site.</td></tr>\n            </tbody>\n          </table>\n          <p>If any check failed, the verdict ends with <strong>Biggest issue</strong>: the summary of the failed check that carries the most weight. It's the single best place to start. If nothing failed, there's no biggest issue, and your work is in the warnings.</p>\n          <p>One thing to know before you share the number: the free scan doesn't run an AI agent through a task, so Task completion shows as \"n/a\" and the score is worked out over Access, Readability and Navigability. A 72 means 72% of the points that could be tested. <a href=\"https://ghostagentlab.com/blog/how-agentscore-works/\">How AgentScore works</a> has the arithmetic.</p>\n          <h2>Then the categories</h2>\n          <p>Each category gets a bar and its points, such as Access 20/25. The app shows the same thing as a share of the category's points, so Access 20/25 reads as 80/100. Use the bars to see where points are being lost, not to rank your work: a low Navigability bar with one failure in it can matter less than a single Access failure.</p>\n          <p>Access comes first for a reason. If AI agents are blocked by your bot protection or robots.txt, they never see the pages the other two categories are about.</p>\n          <h2>What each status means</h2>\n          <table>\n            <thead><tr><th>Status</th><th>What it means</th><th>Counts toward the score</th></tr></thead>\n            <tbody>\n              <tr><td>Passed</td><td>Agents will be fine here.</td><td>Full weight</td></tr>\n              <tr><td>Warning</td><td>It works, but agents will stumble, or a newer standard is missing.</td><td>Half weight</td></tr>\n              <tr><td>Failed</td><td>This stops or seriously misleads agents.</td><td>Nothing</td></tr>\n              <tr><td>Not tested</td><td>The check couldn't run on your site.</td><td>Left out entirely</td></tr>\n            </tbody>\n          </table>\n          <p>Checks are listed failures first, then warnings, then passes, then not tested. Each one has a one-line summary of what was found, often with an example, such as how many buttons have no name or which assistant was blocked.</p>\n          <h2>Don't skip \"not tested\"</h2>\n          <p>A \"not tested\" check never costs you points, and its summary always says why it was skipped. The reasons fall into a few groups:</p>\n          <ul>\n            <li><strong>It doesn't apply.</strong> Guest checkout, checkout fields, product data and agentic commerce are for stores. If the scan found no products or cart, they're skipped.</li>\n            <li><strong>It needs a full cart.</strong> The scanner never adds anything to a cart, so a checkout that only opens with items in it can't be inspected.</li>\n            <li><strong>A page couldn't be rendered.</strong> If the browser couldn't load a page, the checks that depend on it are skipped and say so. If it happens on every scan, ask a developer to check whether that page loads for a first-time desktop visitor.</li>\n            <li><strong>It isn't in the free scan.</strong> \"An AI agent completes a real task\" is always not tested there.</li>\n          </ul>\n          <p>Read these anyway. If you run a store and the report says no product page was found from the home page or sitemap, that's worth knowing in itself: agents look for products the same way the scanner does.</p>\n          <h2>Fixes, severity and score impact</h2>\n          <p>Every warning and failure comes with a fix: the specific change to make. In the free report, you see the findings straight away and unlock the fixes with your email address. In the app, every fix is shown.</p>\n          <p>The app's <strong>Findings</strong> page adds three things that make the list easier to work through:</p>\n          <ul>\n            <li><strong>Severity.</strong> A failure on a check with weight 4 or more is Critical, weight 2 or 3 is High, and weight 1 is Medium. A warning is Medium on a check with weight 3 or more, otherwise Low. A missing llms.txt is labelled Opportunity: worth doing, but a new standard, weighted lightly.</li>\n            <li><strong>Score impact.</strong> How many points your overall score would gain if that finding were fixed. On a store where every Access check was tested, a failed <strong>Bot protection lets AI agents through</strong> is worth about 7 points; a missing llms.txt is worth about 1. The numbers depend on which checks could be tested on your site, so read them from your own report.</li>\n            <li><strong>Why it matters and who usually fixes it.</strong> Open \"Why it matters and how to fix it\" under any finding for a plain-English reason, the fix, and the role that usually owns it: Developer, Bot protection or CDN admin, SEO or content team, or E-commerce platform admin.</li>\n          </ul>\n          <p>The Readiness page shows the top four as <strong>Biggest wins</strong>. The Findings page lets you filter by severity, and <strong>Export CSV</strong> downloads every finding with its severity, score impact, summary, why it matters, fix, who usually fixes it and the scan date, ready for a ticketing tool.</p>\n          <h2>How to decide what to fix first</h2>\n          <ol>\n            <li><strong>Fix Access failures first.</strong> Blocked assistants, challenges on arrival and CAPTCHAs at the checkout stop agents before anything else matters. These usually sit with whoever runs your CDN or bot protection, and are often a settings change rather than a project. Our guide to <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and AI agents</a> covers the common causes.</li>\n            <li><strong>Then the remaining Critical and High findings, by score impact.</strong> Content that only appears after JavaScript, a banner covering the add-to-cart button, and product pages with no price data are typical.</li>\n            <li><strong>Group by owner, not by category.</strong> Sort the CSV by \"Usually fixed by\" and send each person their own short list. One ticket per owner moves faster than one long list for everyone.</li>\n            <li><strong>Pick up cheap warnings along the way.</strong> Image alt text, page landmarks and a missing meta description are small changes that often ship with other work.</li>\n            <li><strong>Leave opportunities until the basics pass.</strong> llms.txt, MCP and agentic commerce are worth doing, and score lightly for a reason: they help most once agents can already get in and read your pages.</li>\n            <li><strong>Rescan after each batch.</strong> In the app you can rescan whenever you like and watch the trend. The free page reuses a domain's result for 24 hours.</li>\n          </ol>\n          <div>\n            <p><strong>A rule of thumb:</strong> if a finding means agents can't get in or can't see your prices, it's this sprint. If it means they'll find it a little harder, it's this quarter. If it's a new standard, it's on the roadmap. Our <a href=\"https://ghostagentlab.com/blog/90-day-agent-readiness-plan/\">90-day plan</a> lays that out week by week.</p>\n          </div>\n          <h2>What the report can't tell you</h2>\n          <p>A high score means your site removes the common obstacles. It doesn't prove that an AI agent can find a product, choose a size and reach the payment step. That's what <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> test in the app: real AI agents run your key journeys, stop at the payment form, and give you a step-by-step replay of where they got stuck. Use the report to clear the path, and <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">journey tests</a> to prove it's clear.</p>",
      "date_published": "2026-10-09T13:16:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Product"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/agent-traffic-is-not-bot-traffic/",
      "url": "https://ghostagentlab.com/blog/agent-traffic-is-not-bot-traffic/",
      "title": "Agent traffic isn't bot traffic",
      "summary": "Lumping AI agents in with scrapers hides customers. Why agent traffic needs its own categories, rules and reports, and how to start separating it.",
      "content_html": "<p>Most websites sort their visitors into two piles: people and bots. People are customers. Bots are a cost, a risk, or noise to filter out of the reports. AI agents don't fit that split. Some of them are the closest thing your site has to a customer in the room, and treating them like scrapers means turning those customers away without knowing it.</p>\n          <h2>The two-pile habit</h2>\n          <p>The people-or-bots split made sense for a long time. Apart from search engine crawlers, which everyone learned to welcome, automated traffic was mostly scrapers, credential stuffers, inventory hoarders and uptime monitors. So the tools grew up around it. Bot protection scores each request on how human it looks. Analytics filters out known bots. Dashboards show \"bot traffic\" as one line, usually as a problem.</p>\n          <p>AI agents arrive in that system looking like bots, because technically they are. But \"bot\" now covers software doing very different jobs for very different reasons, and some of those jobs are done for a specific person who wants to buy something.</p>\n          <h2>Five kinds of automated visitor</h2>\n          <table>\n            <thead><tr><th>Kind</th><th>Examples</th><th>Why it's there</th><th>What it's worth to you</th></tr></thead>\n            <tbody>\n              <tr><td>AI assistants</td><td>ChatGPT-User, Claude-User, Perplexity-User</td><td>A person asked a question just now, and the assistant is reading your page to answer it</td><td>A live customer question about you</td></tr>\n              <tr><td>Browser agents</td><td>Agents that run a real browser to complete a task, such as Ghost Agents</td><td>A person asked it to do something: compare, book, buy</td><td>A customer, part-way through a journey</td></tr>\n              <tr><td>Search and AI search crawlers</td><td>OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot</td><td>Indexing pages so they can appear in search results and AI answers later</td><td>Visibility in tomorrow's answers</td></tr>\n              <tr><td>AI training crawlers</td><td>GPTBot, ClaudeBot, CCBot, Google-Extended</td><td>Collecting content to train models</td><td>A business decision, not a customer</td></tr>\n              <tr><td>Other bots</td><td>SEO tools, uptime monitors, link previews, scrapers</td><td>Their own reasons</td><td>Mostly a cost, sometimes useful</td></tr>\n            </tbody>\n          </table>\n          <p>Google-Extended is a robots.txt token rather than a separate crawler, but it controls the same choice. Running through all five rows is a sixth problem: impostors. Scrapers often claim to be Googlebot or GPTBot to get past bot protection, so a name in a user agent proves nothing on its own.</p>\n          <p>The top two rows are the ones the two-pile habit hurts most. An assistant fetch is one person's question. If it's blocked, that person gets an answer about somebody else. A browser agent that hits a CAPTCHA at checkout is an abandoned cart with no record of why. Neither shows up in your analytics, because most agents never run the tracking script, and both look like \"bots\" to a system tuned to stop bots.</p>\n          <h2>What goes wrong when it's all one pile</h2>\n          <ul>\n            <li><strong>Blocking the wrong visitors.</strong> A rule written to stop scrapers stops ChatGPT-User too, and nobody notices because nobody is measuring it. Our guide to <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and AI agents</a> covers how this happens.</li>\n            <li><strong>Making the wrong call on \"AI\".</strong> Blocking training is a reasonable choice. Blocking every AI user agent to achieve it also removes you from assistants and AI search. <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">Should you block AI crawlers?</a> splits the decision properly.</li>\n            <li><strong>Reading growth as a threat.</strong> \"Bot traffic is up\" sounds like an attack. \"Assistant fetches of our product pages are up\" sounds like demand. They can be the same line on the same chart.</li>\n            <li><strong>Missing the failures.</strong> If agent traffic is filtered out of reporting, so are the 403s, challenge pages and missing pages agents keep hitting. The journey breaks and the dashboard stays green.</li>\n          </ul>\n          <h2>Treat agent traffic as its own channel</h2>\n          <p>The fix isn't to welcome every bot. It's to stop treating them as one thing. In practice that means four habits:</p>\n          <ol>\n            <li><strong>Measure it separately.</strong> Your server or CDN logs see every request, including the ones that never run JavaScript. Split them into the kinds above. <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">How to measure agent traffic</a> shows how.</li>\n            <li><strong>Verify before you trust.</strong> Check a claimed agent against its operator's published IP ranges or reverse DNS, or a request signature where the operator signs its requests with Web Bot Auth. Our guides to <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">verifying AI crawlers</a> and <a href=\"https://ghostagentlab.com/blog/signed-requests/\">signed requests</a> cover the methods.</li>\n            <li><strong>Set rules per kind, not per \"bot\".</strong> Let verified assistants and AI search through to product, pricing, cart and checkout pages. Decide on training crawlers as a separate business question. Rate-limit and challenge the rest, and block impostors outright. <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">Rate limits for AI agents</a> covers the middle ground.</li>\n            <li><strong>Watch outcomes, not just volume.</strong> For each kind, track what it got back: pages served, redirects, blocks and missing pages. A rise in blocked assistant requests is a lost-sales signal, and should reach the same people as a broken checkout.</li>\n          </ol>\n          <div>\n            <p><strong>One honest caveat:</strong> not every agent can be identified. Some browser agents use an ordinary browser's user agent and run from cloud addresses, so they look like people, or like a bot your protection doesn't recognize. Signed requests will help as more operators adopt them. Until then, the best evidence of how these agents fare is to <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">run AI agents through your journeys</a> yourself.</p>\n          </div>\n          <h2>What it means for your reports</h2>\n          <p>The most useful change is often the simplest: stop putting AI agents in the \"bots\" line. Give assistants, browser agents, AI search and training crawlers their own lines, alongside people, and report what each one got back. When a leader asks \"is AI a threat or an opportunity for us?\", that report answers with your own numbers instead of an opinion. Our guide to <a href=\"https://ghostagentlab.com/articles/agent-readiness-kpis/\">agent readiness KPIs</a> suggests what to put in front of them each month, and <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">tracking AI referral traffic</a> covers the visits and sales that follow.</p>\n          <p>In the Ghost Agent Labs app, <strong>Agent traffic</strong> does this split from your server or CDN logs: people, search crawlers, AI training crawlers, AI assistants and autonomous agents, with each assistant's fetches grouped into sessions and a breakdown of how often agents were served, redirected, blocked or sent to missing pages. <strong>Verification</strong> separates verified agents from impostors.</p>\n          <p>Your next customer may arrive as a request with a user agent you've never looked at. It's worth knowing which pile you've put it in.</p>",
      "date_published": "2026-10-09T13:15:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Perspective"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/cdn-settings-ai-agents/",
      "url": "https://ghostagentlab.com/articles/cdn-settings-ai-agents/",
      "title": "Setting up your CDN for AI agents",
      "summary": "How to set up your CDN and WAF for AI agents: verified bot categories, signed requests, rule order, challenge versus block, and logging.",
      "content_html": "<p>Your CDN and web application firewall (WAF) decide, request by request, who gets to see your site. Most were tuned before AI agents were a real source of customers. This guide walks through the settings that matter, in the order your CDN applies them, so verified AI agents get your pages and the bots you don't want still get stopped.</p>\n          <p>If you're still working out why agents get caught in the first place, start with <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">Bot protection and CAPTCHAs: stop blocking the agents you want</a>. This article is the practical follow-on: how to set the CDN up, and how to check it stays that way.</p>\n          <h2>Start with a decision, not a setting</h2>\n          <p>Before anyone changes a rule, agree which kinds of agent you want. Most businesses land somewhere like this:</p>\n          <table>\n            <thead><tr><th>Kind of agent</th><th>What it does</th><th>Typical choice</th></tr></thead>\n            <tbody>\n              <tr><td>AI assistants (for example ChatGPT-User, Claude-User, Perplexity-User)</td><td>Fetch a page because a person asked a question about it right now</td><td>Allow</td></tr>\n              <tr><td>AI search crawlers (for example OAI-SearchBot)</td><td>Index pages so they can appear in AI search answers</td><td>Allow</td></tr>\n              <tr><td>Search engines (Googlebot, Bingbot)</td><td>Classic search indexing, which also feeds some AI features</td><td>Allow</td></tr>\n              <tr><td>AI training crawlers (for example GPTBot, ClaudeBot)</td><td>Collect content to train models</td><td>A business choice; blocking is common</td></tr>\n              <tr><td>Unverified bots claiming any of the names above</td><td>Unknown</td><td>Challenge, rate limit or block</td></tr>\n            </tbody>\n          </table>\n          <p>Write the decision down. It's what your CDN admin implements, and it's what you check against later. Our guides to <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a> and <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">whether to block or allow AI crawlers</a> cover the trade-offs.</p>\n          <h2>Use verified bot categories, not user agent strings</h2>\n          <p>Most large CDNs and bot management products keep a list of verified bots: crawlers and agents whose requests they have confirmed really come from the operator they name. Many also group those bots into categories, so you can write one rule for \"AI assistants\" instead of maintaining a list of names yourself.</p>\n          <p>Build your rules on that verification, not on the <code>User-Agent</code> header. Anyone can send a request that says it's ChatGPT-User. A rule that skips bot checks for that string is an open door for scrapers. Our guide to <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">telling real AI crawlers from fakes</a> explains the methods CDNs use behind the scenes: published IP ranges, reverse DNS and signatures.</p>\n          <div>\n            <p><strong>Check the category names in your own console.</strong> Vendors name and split these categories differently, and they change them as new agents appear. Confirm which category each agent you care about sits in before you write a rule that depends on it.</p>\n          </div>\n          <h2>Turn on signature verification where you can</h2>\n          <p>IP-based verification struggles with agents that run in the cloud, because their addresses change and are often shared. Web Bot Auth fixes that by having the agent sign each request with a private key, and publish the matching public key on its own domain. It builds on HTTP Message Signatures (<a href=\"https://www.rfc-editor.org/rfc/rfc9421\">RFC 9421</a>), and is still being standardized, so support varies.</p>\n          <p>If your CDN can verify these signatures, turn it on and treat a valid signature as verification. If it can't yet, nothing breaks: signed requests are ordinary requests with extra headers. Our post <a href=\"https://ghostagentlab.com/blog/signed-requests/\">Why we sign every request our agents send</a> explains how it works in practice.</p>\n          <h2>Get the rule order right</h2>\n          <p>CDNs apply rules in a sequence, and the first matching action often wins. Most agent problems we see come from order, not from a missing rule. A sensible order looks like this:</p>\n          <ol>\n            <li><strong>Security rules for everyone.</strong> Managed WAF rules that catch attacks such as SQL injection should apply to every request, verified bots included. Verification proves who sent a request, not that the request is safe.</li>\n            <li><strong>Hard blocks you need for legal or business reasons.</strong> For example, blocked training crawlers, or country restrictions you're required to enforce. See <a href=\"https://ghostagentlab.com/articles/geo-blocking-ai-agents/\">Geo-blocking, VPN blocks and AI agents</a> for the country case.</li>\n            <li><strong>Allow verified agents you want</strong>, by category or verified identity. \"Allow\" here should mean \"skip bot challenges and bot scoring\", not \"skip all security\".</li>\n            <li><strong>Rate limits</strong> for everyone, including verified agents, set high enough for normal use. See <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a>.</li>\n            <li><strong>Bot scoring and challenges</strong> for whatever is left: unverified automation, impostors and suspicious traffic.</li>\n          </ol>\n          <p>The common mistake is a broad \"block AI bots\" or \"challenge automated traffic\" rule placed above the verified-agent allow rule. The allow rule then never runs.</p>\n          <h2>Challenge or block: choose deliberately</h2>\n          <p>Most bot tools offer several actions. For a person they feel very different. For an AI agent, most of them end the same way.</p>\n          <table>\n            <thead><tr><th>Action</th><th>What a person sees</th><th>What an AI agent gets</th></tr></thead>\n            <tbody>\n              <tr><td>Block</td><td>An error page</td><td>An error page; it tells the user it can't reach you</td></tr>\n              <tr><td>JavaScript or \"managed\" challenge</td><td>A short wait, usually invisible</td><td>Usually a dead end: most assistant fetchers don't run scripts</td></tr>\n              <tr><td>Interactive challenge or CAPTCHA</td><td>A puzzle or checkbox</td><td>A dead end: agents can't, and shouldn't, solve them</td></tr>\n              <tr><td>Rate limit (429 with Retry-After)</td><td>Rarely seen</td><td>A clear signal to slow down and try again</td></tr>\n              <tr><td>Log only</td><td>Nothing</td><td>The page</td></tr>\n            </tbody>\n          </table>\n          <p>So treat a challenge as a block when you're thinking about agents. Keep challenges for traffic you genuinely don't want, and for high-risk actions such as login, account creation and payment. Product, pricing, policy and help pages should open for verified agents without one.</p>\n          <h2>Log enough to see what happened</h2>\n          <p>When an assistant says it \"couldn't access\" your site, someone has to find out why. Make sure your CDN logs, or the samples it keeps, include for each request:</p>\n          <ul>\n            <li>The user agent, and whether the CDN treated the request as a verified bot (and in which category)</li>\n            <li>The action taken (allowed, challenged, blocked, rate limited) and the ID of the rule that took it</li>\n            <li>The response status code and the path</li>\n            <li>The client country and network, if you use location rules</li>\n          </ul>\n          <p>Before you switch a new rule to block, run it in log-only (sometimes called \"simulate\" or \"count\") mode for a few days and check what it would have caught. Then review blocked and challenged agent traffic regularly. Our guides to <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">measuring AI agent traffic</a> and <a href=\"https://ghostagentlab.com/articles/agent-errors-in-logs/\">finding the errors AI agents hit in your logs</a> show what to look for.</p>\n          <h2>Notes for common vendors</h2>\n          <p>Products change quickly and menus move, so treat these as pointers, and check your vendor's current documentation for the exact names and the plan you're on.</p>\n          <ul>\n            <li><strong>Cloudflare</strong> maintains a verified bots program and documents bot categories that you can reference in custom rules, including separate groupings for AI crawlers and other AI agents. It also offers settings aimed at blocking AI crawlers; check exactly which agents those cover before turning them on. Cloudflare has published work on Web Bot Auth, so check whether signature verification is available on your plan.</li>\n            <li><strong>Akamai</strong> bot management sorts known bots into categories and lets you set an action for each. Find where the AI-related categories sit and which action they currently get.</li>\n            <li><strong>Fastly</strong> offers bot management and WAF features with its own way of identifying known bots. Check how verified bots are flagged in its rules, and where that check runs relative to your other rules.</li>\n            <li><strong>AWS WAF</strong> has a bot control managed rule group that labels requests, including whether a bot is verified. Rules later in the web ACL can act on those labels, so order and labels together decide what agents get.</li>\n          </ul>\n          <p>If you use a dedicated bot product in front of or behind your CDN, the same questions apply to it. Two layers mean two places an agent can be stopped.</p>\n          <h2>Check it from the outside</h2>\n          <p>A configuration that looks right in the console can still behave differently in practice. Test it as an agent would:</p>\n          <div>\n            <ol>\n              <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>. It requests your home page as a normal browser and as AI agents including ChatGPT-User, Claude-User, Perplexity-User and OAI-SearchBot, and fails the bot protection check if an agent is blocked or challenged while the browser gets through. Other checks test your product, pricing and cart pages the same way, and look for CAPTCHAs on arrival and at checkout.</li>\n              <li>Look at your CDN's logs for the scan and confirm which rule acted on each request.</li>\n              <li>Re-run the scan after every bot or WAF change, and after your vendor announces changes to its bot categories.</li>\n            </ol>\n          </div>\n          <p>One caution: AgentScore's requests under those agent names aren't signed and don't come from those operators' networks, so a correctly configured CDN that only allows <em>verified</em> agents may still challenge them. That's the right behavior. Read a failure alongside your logs: if real, verified agent traffic is getting through, your setup is working. Our post on <a href=\"https://ghostagentlab.com/blog/scanner-wears-other-names/\">why our scanner wears other names</a> explains the trade-off.</p>",
      "date_published": "2026-10-09T13:14:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/geo-blocking-ai-agents/",
      "url": "https://ghostagentlab.com/articles/geo-blocking-ai-agents/",
      "title": "Geo-blocking, VPN blocks and AI agents",
      "summary": "Why country, VPN and data center blocks and IP-based redirects catch AI agents visiting from cloud servers abroad, and how to let verified agents through.",
      "content_html": "<p>A shopper in Manchester asks an AI assistant about your store. The assistant doesn't visit from Manchester. It visits from a cloud data center, quite possibly in another country. If your site blocks, challenges or redirects visitors by location or network, that one fact can decide whether the assistant sees your real store, the wrong store, or nothing at all.</p>\n          <h2>Where AI agents actually come from</h2>\n          <p>When a person asks an assistant a question, the page request doesn't come from their phone or laptop. It comes from the assistant's servers. The same is true for AI search crawlers and for browser agents that run in the cloud. That has three consequences:</p>\n          <ul>\n            <li><strong>The country is the server's, not the shopper's.</strong> Agent traffic is concentrated in the regions where AI companies run their infrastructure. Those are rarely the same as your customers' locations.</li>\n            <li><strong>The network is a data center.</strong> Agent requests come from cloud and hosting providers, the same networks that a lot of abusive automation uses. Reputation lists often score them as high risk.</li>\n            <li><strong>Addresses change.</strong> Agents that run on shared cloud infrastructure don't always come from fixed, published IP ranges.</li>\n          </ul>\n          <p>None of this is a sign of bad intent. It's simply how AI agents work. But it means rules written with human visitors in mind can treat agents very differently from the people they're acting for.</p>\n          <h2>Four rules that catch agents</h2>\n          <h3>Country blocks</h3>\n          <p>Some sites block whole countries, often to cut fraud or because they don't ship there. If an assistant's servers sit in a blocked country, every request it makes for every shopper is turned away, including shoppers in countries you serve.</p>\n          <h3>VPN, proxy and data center blocks</h3>\n          <p>Blocking \"anonymizing\" networks or hosting providers is a common anti-fraud and anti-scraping setting. It catches nearly all AI agent traffic, because nearly all of it comes from hosting providers. It's one of the most common reasons an assistant can reach a competitor's site but not yours.</p>\n          <h3>Automatic geo-redirects</h3>\n          <p>Many international stores redirect visitors to a country site based on IP address. An agent fetching your UK product page for a UK shopper may be sent to your US site instead, then quote US prices, US stock and US delivery times. The answer looks confident and is wrong. Some redirects go to a country picker page, which leaves the agent with no product at all.</p>\n          <h3>Location-based content</h3>\n          <p>Some sites quietly change prices, currency, product ranges or policies based on the visitor's location, without changing the URL. An agent then describes a version of your store the shopper will never see.</p>\n          <h2>What to do instead</h2>\n          <p>The aim is to keep the protection you need while letting verified agents see the right content.</p>\n          <h3>1. Separate legal requirements from preferences</h3>\n          <p>Some location rules are required: sanctions, licensing, or content you're not allowed to show in certain places. Keep those, and get advice on how they apply to automated visitors. Many others are preferences, such as \"we don't ship there, so block it\". For those, ask whether blocking the server's location actually achieves anything. You already check the delivery address at checkout, which is where a shipping restriction belongs.</p>\n          <h3>2. Let verified agents past location and network rules</h3>\n          <p>Most CDNs and bot management tools can tell verified bots apart from other traffic. Add an exception so that verified AI assistants and AI search agents skip the country, VPN and data center rules that are about fraud or scraping. Put it in the right place in your rule order, and never base it on the user agent alone. Our guide to <a href=\"https://ghostagentlab.com/articles/cdn-settings-ai-agents/\">setting up your CDN for AI agents</a> covers rule order, and <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">telling real AI crawlers from fakes</a> covers verification.</p>\n          <div>\n            <p><strong>Don't allow whole cloud providers.</strong> Lifting a data center block for an entire cloud network lets in every scraper hosted there. Allow verified agents, and keep the block for everything else.</p>\n          </div>\n          <h3>3. Let the URL decide the country, not the IP address</h3>\n          <p>Give each country or language its own URL, such as <code>northwind.example/en-gb/</code> or <code>uk.northwind.example</code>, and serve the content that URL promises, whoever asks. If you want to steer people to their local site, show a banner suggesting it instead of redirecting. Agents and people can then follow the link that matches the shopper. Mark the alternatives with <code>hreflang</code> so agents and search engines can find them:</p>\n          <pre><code>&lt;link rel=\"alternate\" hreflang=\"en-gb\" href=\"https://northwind.example/en-gb/boots/trail-runner/\"&gt;\n&lt;link rel=\"alternate\" hreflang=\"en-us\" href=\"https://northwind.example/en-us/boots/trail-runner/\"&gt;\n&lt;link rel=\"alternate\" hreflang=\"x-default\" href=\"https://northwind.example/boots/trail-runner/\"&gt;</code></pre>\n          <p>Our guide to <a href=\"https://ghostagentlab.com/articles/international-sites-ai-agents/\">languages, currencies and regions</a> goes further on international setups.</p>\n          <h3>4. Make the default page useful</h3>\n          <p>Some agents will still land on your default site. Make sure it's a real store, not a country picker. State which country and currency the prices are in, and link clearly to the other country sites. Put the same information in your structured data, so an agent knows the price it read is in US dollars and not pounds. See <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">machine-readable prices</a>.</p>\n          <h3>5. Use rate limits for the rest</h3>\n          <p>For traffic you can't verify, from data centers or anywhere else, rate limits are usually a better tool than outright blocks. They stop bulk scraping while still letting through a small number of requests made on someone's behalf. See <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a>.</p>\n          <h2>How to find out if this affects you</h2>\n          <table>\n            <thead><tr><th>Look for</th><th>Where</th><th>What it suggests</th></tr></thead>\n            <tbody>\n              <tr><td>Country, ASN, \"hosting\" or \"anonymous proxy\" rules</td><td>CDN, WAF or bot management settings</td><td>Agents may be blocked or challenged by location or network</td></tr>\n              <tr><td>Redirects based on IP country</td><td>CDN rules, e-commerce platform settings, geolocation apps or plugins</td><td>Agents may see the wrong country's store</td></tr>\n              <tr><td>403s, challenges or 302 redirects for ChatGPT-User, Claude-User or Perplexity-User</td><td>CDN or server logs</td><td>Real agent visits are being turned away or sent elsewhere</td></tr>\n              <tr><td>Assistants quoting the wrong currency or prices</td><td>Ask a few assistants about your products</td><td>A redirect or location-based content is reaching agents</td></tr>\n            </tbody>\n          </table>\n          <p>An <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> scan helps too. Its requests come from cloud servers, like real agents' requests do, and it compares what AI agents get from your home page, product, pricing and cart pages with what a normal browser gets from the same servers. It can't show what visitors in every country see, so a block that only applies to certain countries may not show up. Pair it with a look at your rules and logs. Our guide to <a href=\"https://ghostagentlab.com/articles/agent-errors-in-logs/\">finding the errors AI agents hit in your logs</a> shows how to pull out agent requests and their status codes.</p>\n          <h2>A short checklist</h2>\n          <div>\n            <ol>\n              <li>List every location and network rule on your CDN, WAF, bot tool and e-commerce platform, and note which are legally required.</li>\n              <li>Add an exception for verified AI assistants and AI search agents to the rules that aren't.</li>\n              <li>Replace automatic IP redirects with country URLs, <code>hreflang</code> and a suggestion banner.</li>\n              <li>Make sure your default site shows real products, with the currency and country stated.</li>\n              <li>Check your logs for agent requests that get 403s, challenges or redirects, and re-run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> after each change.</li>\n            </ol>\n          </div>",
      "date_published": "2026-10-09T13:13:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/login-walls-ai-agents/",
      "url": "https://ghostagentlab.com/articles/login-walls-ai-agents/",
      "title": "Login walls and gated content: what AI agents can reach",
      "summary": "What AI agents can't reach behind logins and paywalls, what to keep public, how to hand off to the customer, and paywall structured data for publishers.",
      "content_html": "<p>An AI agent acting for a customer usually doesn't have that customer's password, and it can't create accounts or pay for subscriptions on a whim. Anything you put behind a login or a paywall is, for most agents, simply not there. That's often the right call. But many sites hide things that should be public, and lose recommendations because of it.</p>\n          <h2>What agents can and can't reach</h2>\n          <p>Different kinds of agent meet a login wall in different ways:</p>\n          <table>\n            <thead><tr><th>Kind of agent</th><th>Behind a login or paywall</th></tr></thead>\n            <tbody>\n              <tr><td>AI assistants fetching a page for a question</td><td>Get the login page or the paywall, and answer from whatever is visible</td></tr>\n              <tr><td>AI search and training crawlers</td><td>Index only what's public, unless you choose to show them more</td></tr>\n              <tr><td>Browser agents carrying out a task</td><td>Usually stop and ask the person to sign in, or give up. Some can hand control back to the person for the sign-in step</td></tr>\n            </tbody>\n          </table>\n          <p>So the question for each page is simple: if an agent can only see the public version, can it still answer the customer's question and move them forward?</p>\n          <h2>What to keep public</h2>\n          <p>These are the things agents are most often asked about. If they sit behind a login, the agent either says it doesn't know or quotes a competitor who shows them openly.</p>\n          <ul>\n            <li><strong>Prices.</strong> \"Sign in to see price\" and \"add to cart to see price\" hide the one fact shoppers ask about most. If you can't show an exact price, show a starting price or a range. See <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">machine-readable prices</a>.</li>\n            <li><strong>Product information.</strong> Specifications, sizes, materials, compatibility and stock. A catalog that only appears after sign-in is invisible to agents.</li>\n            <li><strong>Policies.</strong> Shipping, returns, warranty and payment options. These decide many purchases, and agents quote them directly. See <a href=\"https://ghostagentlab.com/articles/policy-pages-ai-assistants/\">shipping, returns and FAQ pages AI assistants can quote</a>.</li>\n            <li><strong>Plans and pricing for software.</strong> What each plan includes and costs, and whether there's a free trial. See <a href=\"https://ghostagentlab.com/blog/agent-readiness-for-saas/\">agent readiness for software companies</a>.</li>\n            <li><strong>Help content.</strong> Setup guides, FAQs and documentation. If your help center needs a login, assistants can't answer support questions about your product.</li>\n            <li><strong>Contact and location details.</strong> Opening hours, addresses and how to reach you.</li>\n          </ul>\n          <p>Trade and wholesale businesses often hide everything, because prices vary by customer. Consider showing the catalog and list prices publicly, with a note that account holders get their own prices after signing in. The agent can then recommend you and send the buyer to log in for the final number.</p>\n          <h2>What can stay behind the wall</h2>\n          <p>Plenty of content belongs there: order history, saved addresses, account settings, customer-specific prices, members-only content and paid articles. The goal isn't to remove logins. It's to make sure the wall sits in front of private and paid things, not in front of the information a customer needs to decide.</p>\n          <h2>Handing off to the person</h2>\n          <p>When an agent does reach something that needs the customer, make the handoff smooth:</p>\n          <ul>\n            <li><strong>Say what's needed in plain text.</strong> \"Sign in to see your order history\" is something an agent can relay. A login form with no explanation isn't.</li>\n            <li><strong>Return to the same place.</strong> After sign-in, send the person back to the page or cart they were on, with the cart intact, so the agent's work isn't lost.</li>\n            <li><strong>Don't require an account to buy.</strong> Guest checkout lets an agent finish a purchase without creating an account for someone. AgentScore checks for it in <em>Shoppers can check out without an account</em>, and our guide to <a href=\"https://ghostagentlab.com/articles/guest-checkout-ai-agents/\">guest checkout</a> explains why it matters.</li>\n            <li><strong>Use proper status codes.</strong> Redirect to a login page, or return 401 or 403, rather than a 200 page that looks like content but isn't. Agents and crawlers can then tell \"this needs a login\" from \"this page says nothing\".</li>\n            <li><strong>Make sign-in forms agent-friendly.</strong> Labelled fields and standard <code>autocomplete</code> values help browser agents and password managers alike. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-sign-up-booking/\">sign-up, booking and lead forms AI agents can finish</a>.</li>\n          </ul>\n          <h2>Paywalls and structured data for publishers</h2>\n          <p>Publishers face a specific version of this. You may want search engines to see the full article so they can rank it, while readers see a paywall. Showing crawlers different content from readers can look like cloaking. Google's documented answer is paywalled content structured data: you mark the article as not free, and point to the part that's behind the paywall.</p>\n          <pre><code>&lt;script type=\"application/ld+json\"&gt;\n{\n  \"@context\": \"https://schema.org\",\n  \"@type\": \"NewsArticle\",\n  \"headline\": \"How trail running shoes are tested\",\n  \"isAccessibleForFree\": false,\n  \"hasPart\": {\n    \"@type\": \"WebPageElement\",\n    \"isAccessibleForFree\": false,\n    \"cssSelector\": \".paywalled-content\"\n  }\n}\n&lt;/script&gt;</code></pre>\n          <p>The <code>cssSelector</code> must match a class on the element that wraps the paid content. Google explains the details in its <a href=\"https://developers.google.com/search/docs/appearance/structured-data/paywalled-content\">paywalled content guidelines</a>, and the properties are defined at <a href=\"https://schema.org/isAccessibleForFree\">schema.org</a>.</p>\n          <p>Two things to keep in mind:</p>\n          <ul>\n            <li><strong>The markup describes access; it doesn't grant it.</strong> Whether you show full articles to any crawler, including AI search and training crawlers, is a business and licensing decision. Make it deliberately, per agent, and reflect it in <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt</a> and your CDN rules.</li>\n            <li><strong>Give agents something to cite.</strong> A clear headline, a summary paragraph and the key facts above the paywall let assistants describe and link to your article even when they can't read it all. A paywall that hides everything, headline included, gets you nothing.</li>\n          </ul>\n          <h2>How to check your own site</h2>\n          <div>\n            <ol>\n              <li>Open a private browser window, signed out, and try to find a product's price, your returns policy and your contact details. Anything you can't find, an agent can't either.</li>\n              <li>Add an item to the cart and go to checkout. If you're asked to sign in or create an account with no guest option, agents buying for someone will stop there.</li>\n              <li>Ask a few AI assistants what your products cost and what your returns policy is. Wrong or missing answers often point to content behind a wall.</li>\n              <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> for the guest checkout, key pages and price checks, and use <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> to run a real journey and see exactly where an agent gets stuck.</li>\n            </ol>\n          </div>",
      "date_published": "2026-10-09T13:12:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/organization-structured-data/",
      "url": "https://ghostagentlab.com/articles/organization-structured-data/",
      "title": "Business and store details AI agents can trust: Organization and LocalBusiness data",
      "summary": "How to describe your business, contact points, opening hours and return policy with Organization and LocalBusiness JSON-LD that AI agents can trust.",
      "content_html": "<p>Before an AI agent recommends a store, it wants to know who it's dealing with: the business name, how to contact it, where it ships from, when it's open, and what happens if something goes back. Organization data puts those facts in one place, in a form an agent can read without piecing them together from your footer.</p>\n          <h2>Why business details matter to AI agents</h2>\n          <p>People ask AI assistants questions like \"is this a real shop?\", \"do they have a phone number?\", \"are they open on Sunday?\" and \"can I return it if it doesn't fit?\". To answer, an agent has to find those facts somewhere on your site. Usually they're scattered: the phone number in the footer, the address on a contact page, the returns window in a policy page, the opening hours in an image.</p>\n          <p>An agent can often work it out, but every guess is a chance to get it wrong or to give up and recommend someone whose details are easier to confirm. Structured data about your business gives it one trusted summary, on the page it's most likely to visit first: your home page.</p>\n          <p>Product markup describes what you sell. Organization markup describes who is selling it. Agents use both. For products, see <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data that AI shopping agents can read</a>.</p>\n          <h2>Pick the right type</h2>\n          <p>Schema.org has a family of types for businesses. Use the most specific one that's true.</p>\n          <table>\n            <thead><tr><th>Type</th><th>Use it when</th></tr></thead>\n            <tbody>\n              <tr><td><code>Organization</code></td><td>The general case: any company, brand or nonprofit</td></tr>\n              <tr><td><code>OnlineStore</code></td><td>You sell online and have no shop customers can walk into. It's a more specific kind of <code>Organization</code></td></tr>\n              <tr><td><code>LocalBusiness</code> or a subtype such as <code>Store</code></td><td>Customers can visit you at a physical address, such as a shop, showroom or clinic</td></tr>\n            </tbody>\n          </table>\n          <p>If you have an online store and physical shops, describe the company as an <code>Organization</code> or <code>OnlineStore</code> on the home page, and give each shop its own <code>LocalBusiness</code> (or <code>Store</code>) markup on its own location page, with its own address and hours.</p>\n          <h2>A complete example</h2>\n          <p>Here's JSON-LD for the home page of a fictional online coffee store. It goes in the page's HTML, inside the <code>&lt;head&gt;</code> or <code>&lt;body&gt;</code>:</p>\n          <pre><code>&lt;script type=\"application/ld+json\"&gt;\n{\n  \"@context\": \"https://schema.org\",\n  \"@type\": \"OnlineStore\",\n  \"@id\": \"https://northwind.example/#organization\",\n  \"name\": \"Northwind Coffee\",\n  \"legalName\": \"Northwind Coffee Ltd\",\n  \"url\": \"https://northwind.example/\",\n  \"logo\": \"https://northwind.example/images/logo.png\",\n  \"description\": \"Specialty coffee roasted to order and shipped across the US.\",\n  \"email\": \"help@northwind.example\",\n  \"telephone\": \"+1-555-010-0199\",\n  \"address\": {\n    \"@type\": \"PostalAddress\",\n    \"streetAddress\": \"120 Harbor Street\",\n    \"addressLocality\": \"Portland\",\n    \"addressRegion\": \"OR\",\n    \"postalCode\": \"97201\",\n    \"addressCountry\": \"US\"\n  },\n  \"contactPoint\": [{\n    \"@type\": \"ContactPoint\",\n    \"contactType\": \"customer service\",\n    \"email\": \"help@northwind.example\",\n    \"telephone\": \"+1-555-010-0199\",\n    \"availableLanguage\": [\"en\"],\n    \"hoursAvailable\": {\n      \"@type\": \"OpeningHoursSpecification\",\n      \"dayOfWeek\": [\"Monday\", \"Tuesday\", \"Wednesday\", \"Thursday\", \"Friday\"],\n      \"opens\": \"09:00\",\n      \"closes\": \"17:00\"\n    }\n  }],\n  \"sameAs\": [\n    \"https://www.instagram.com/northwindcoffee.example\",\n    \"https://www.linkedin.com/company/northwind-coffee-example\"\n  ],\n  \"hasMerchantReturnPolicy\": {\n    \"@type\": \"MerchantReturnPolicy\",\n    \"applicableCountry\": \"US\",\n    \"returnPolicyCategory\": \"https://schema.org/MerchantReturnFiniteReturnWindow\",\n    \"merchantReturnDays\": 30,\n    \"returnMethod\": \"https://schema.org/ReturnByMail\",\n    \"returnFees\": \"https://schema.org/FreeReturn\",\n    \"merchantReturnLink\": \"https://northwind.example/returns\"\n  }\n}\n&lt;/script&gt;</code></pre>\n          <h2>What each part tells an agent</h2>\n          <ul>\n            <li><strong>Name, legal name, URL and logo.</strong> Confirms which business this is, and that this website belongs to it. Use the name customers know you by in <code>name</code>, and your registered company name in <code>legalName</code>.</li>\n            <li><strong>Contact points.</strong> A <code>ContactPoint</code> for each way to reach you, with <code>contactType</code> saying what it's for (\"customer service\", \"sales\", \"technical support\"). Add the hours it's staffed and the languages it covers, so an agent can tell someone whether they'll get an answer today.</li>\n            <li><strong>Address.</strong> A full <code>PostalAddress</code> answers \"where are they based?\" and helps an agent judge shipping times and which country's consumer rules apply. Only include it if you're happy for it to be quoted.</li>\n            <li><strong><code>sameAs</code>.</strong> Links to your official profiles elsewhere: social accounts, marketplace storefronts, a Wikipedia or business-register entry. They help an agent connect your site to what others say about you, and make it harder to confuse you with a similarly named business.</li>\n            <li><strong>Return policy.</strong> <code>hasMerchantReturnPolicy</code> on the organization sets a default for everything you sell. Google added support for return policies at the organization level, so you don't have to repeat it on every product. Product-level policies can still override it where they differ.</li>\n          </ul>\n          <h2>Opening hours for places customers visit</h2>\n          <p>For a shop, showroom or other location, the most common question is \"are they open now?\". Use <code>openingHoursSpecification</code> on the <code>LocalBusiness</code>, one entry per set of days with the same hours:</p>\n          <pre><code>\"openingHoursSpecification\": [\n  { \"@type\": \"OpeningHoursSpecification\",\n    \"dayOfWeek\": [\"Monday\", \"Tuesday\", \"Wednesday\", \"Thursday\", \"Friday\"],\n    \"opens\": \"08:00\", \"closes\": \"18:00\" },\n  { \"@type\": \"OpeningHoursSpecification\",\n    \"dayOfWeek\": \"Saturday\", \"opens\": \"09:00\", \"closes\": \"16:00\" }\n]</code></pre>\n          <p>Leave out days you're closed. For holidays, add an entry with <code>validFrom</code> and <code>validThrough</code> dates, so an agent doesn't send someone to a locked door on December 25. Add <code>geo</code> coordinates and a <code>telephone</code> for each location too.</p>\n          <h2>Rules that keep it trustworthy</h2>\n          <ul>\n            <li><strong>Match what people see.</strong> The phone number, address, hours and return window in your markup must match your contact and policy pages. If they disagree, an agent can't tell which is right, and may say so.</li>\n            <li><strong>One source of truth.</strong> Generate the markup from the same settings as your footer and contact page, so a change in one place updates both. Hand-written markup goes stale the first time your hours change.</li>\n            <li><strong>Put it in the HTML.</strong> Markup added by a tag manager or other JavaScript after the page loads is invisible to agents that don't run JavaScript.</li>\n            <li><strong>Describe it once.</strong> Use a stable <code>@id</code> (like <code>https://northwind.example/#organization</code>) so other markup on your site, such as a product's <code>brand</code> or a <code>WebSite</code>'s <code>publisher</code>, can point to the same organization instead of repeating it.</li>\n            <li><strong>Keep the policy pages too.</strong> Markup summarizes; your returns and shipping pages explain. Agents read both. See <a href=\"https://ghostagentlab.com/articles/policy-pages-ai-assistants/\">policy pages AI assistants can quote</a>.</li>\n          </ul>\n          <h2>What AgentScore checks</h2>\n          <p>The <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> check \"Structured data describes your business and products\" reads your home page's HTML for schema.org JSON-LD:</p>\n          <ul>\n            <li>It <strong>passes</strong> when JSON-LD is in the HTML before any JavaScript runs, and lists the types it found, such as <code>Organization</code> and <code>WebSite</code>.</li>\n            <li>It <strong>warns</strong> if the JSON-LD only appears after JavaScript runs, or if the page only has microdata.</li>\n            <li>It <strong>fails</strong> if there's no schema.org structured data on the home page at all.</li>\n          </ul>\n          <p>The check confirms the markup is there and readable. It doesn't judge whether your hours or phone number are right, so validate the content yourself.</p>\n          <h2>How to add it</h2>\n          <ol>\n            <li><strong>See what you have.</strong> View your home page's source and search for <code>application/ld+json</code>. Many platforms and themes already output a basic <code>Organization</code> block with just a name and logo.</li>\n            <li><strong>Collect the facts.</strong> Ask customer service for the contact details and hours they actually want shared, and check the return window against your current policy.</li>\n            <li><strong>Extend the markup.</strong> Your developer can extend the theme template, or an SEO app or plugin can add the fields. Prefer one that reads from your store settings.</li>\n            <li><strong>Validate.</strong> Run the page through the <a href=\"https://validator.schema.org/\">Schema Markup Validator</a>, and Google's Rich Results Test for the fields Google uses.</li>\n            <li><strong>Re-check after changes.</strong> Add \"update the structured data\" to the checklist for changing hours, phone numbers or policies.</li>\n          </ol>\n          <p>Business details are one part of being readable to AI agents. For the rest, see the <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness guide</a>, and for the details shoppers ask about most, <a href=\"https://ghostagentlab.com/articles/reviews-ratings-ai-agents/\">reviews and ratings AI agents can read</a>.</p>",
      "date_published": "2026-10-09T13:11:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/product-descriptions-ai-agents/",
      "url": "https://ghostagentlab.com/articles/product-descriptions-ai-agents/",
      "title": "Product descriptions AI agents can compare",
      "summary": "How to write product descriptions AI agents can compare: specs as text and tables, consistent attributes, units, materials and compatibility.",
      "content_html": "<p>When someone asks an AI agent for \"a waterproof hiking jacket under 500 grams that packs into its own pocket\", the agent has to read your product pages and compare them with everyone else's. It can only recommend what it can confirm. Product descriptions written as facts, not mood, are what get you onto the shortlist.</p>\n          <h2>How AI agents compare products</h2>\n          <p>A person browsing can look at a photo, skim the copy and get a feel for a product. An AI agent works differently. It turns the shopper's request into requirements (waterproof, under 500 grams, packable, under $200) and then checks each product page for evidence that the product meets them.</p>\n          <p>If your page says \"featherlight protection for every adventure\", the agent can't match that to \"under 500 grams\". If a competitor's page says \"Weight: 410 g (size M)\", it can. The agent recommends the product it can back up, and yours drops out, even if it's the better jacket.</p>\n          <p>This doesn't mean brand voice has to go. It means the facts need to be on the page too, written so they can be read and compared on their own.</p>\n          <h2>Write specs as text, not images</h2>\n          <p>Agents read text. Specifications that live only in an image, a size-chart graphic or a downloadable PDF are often invisible to them. The same goes for details that only appear when someone clicks a tab or accordion, if that content is loaded by JavaScript rather than present in the HTML. (Content hidden by CSS until clicked is fine; it's still in the page.)</p>\n          <p>Put the key facts in the page's HTML as plain text. A simple table works well, because each row is a clear pair of attribute and value:</p>\n          <pre><code>&lt;table&gt;\n  &lt;caption&gt;Specifications&lt;/caption&gt;\n  &lt;tr&gt;&lt;th scope=\"row\"&gt;Weight&lt;/th&gt;&lt;td&gt;410 g (size M)&lt;/td&gt;&lt;/tr&gt;\n  &lt;tr&gt;&lt;th scope=\"row\"&gt;Waterproof rating&lt;/th&gt;&lt;td&gt;20,000 mm hydrostatic head&lt;/td&gt;&lt;/tr&gt;\n  &lt;tr&gt;&lt;th scope=\"row\"&gt;Shell fabric&lt;/th&gt;&lt;td&gt;100% recycled nylon, 3-layer&lt;/td&gt;&lt;/tr&gt;\n  &lt;tr&gt;&lt;th scope=\"row\"&gt;Packed size&lt;/th&gt;&lt;td&gt;18 x 12 x 7 cm, packs into chest pocket&lt;/td&gt;&lt;/tr&gt;\n  &lt;tr&gt;&lt;th scope=\"row\"&gt;Fit&lt;/th&gt;&lt;td&gt;Regular, room for a fleece underneath&lt;/td&gt;&lt;/tr&gt;\n&lt;/table&gt;</code></pre>\n          <p>A definition list (<code>&lt;dl&gt;</code>) works just as well. What matters is that each fact has a label next to it, in text.</p>\n          <h2>Make attributes comparable</h2>\n          <p>Comparison only works when the same attribute is described the same way across your catalog. Agree on a set of attributes for each category and fill them for every product.</p>\n          <table>\n            <thead><tr><th>Do</th><th>Avoid</th></tr></thead>\n            <tbody>\n              <tr><td>Weight: 410 g</td><td>Ultralight</td></tr>\n              <tr><td>Battery life: up to 14 hours of video playback</td><td>All-day battery</td></tr>\n              <tr><td>Capacity: 1.7 liters (7 cups)</td><td>Family size</td></tr>\n              <tr><td>Dimensions: 60 x 40 x 25 cm (W x D x H)</td><td>Compact design</td></tr>\n              <tr><td>Material: 100% organic cotton, 180 gsm</td><td>Premium soft fabric</td></tr>\n              <tr><td>Fits: iPhone 15 and iPhone 16 (not Pro Max)</td><td>Fits most phones</td></tr>\n            </tbody>\n          </table>\n          <p>Use the same attribute names, the same order and the same units on every product in a category. An agent comparing your three kettles should find \"Capacity\" in the same place on each page.</p>\n          <h2>Units, materials and measurements</h2>\n          <ul>\n            <li><strong>Always give the unit.</strong> \"Width: 60\" means nothing. \"Width: 60 cm\" does. Where your customers use both systems, give both: \"60 cm (23.6 in)\".</li>\n            <li><strong>Say what was measured.</strong> \"Weight: 410 g (size M)\" or \"Battery life: up to 14 hours (video playback, 50% brightness)\". Conditions turn a claim into a fact.</li>\n            <li><strong>Name the materials.</strong> Give the fiber, metal or wood, the percentage in blends, and any certification by its proper name. \"Vegan leather\" is a category; \"polyurethane-coated cotton\" is a material.</li>\n            <li><strong>Give exact sizes.</strong> Link to a size guide in HTML text, not only an image, with body measurements per size.</li>\n            <li><strong>State what's in the box.</strong> Agents get asked \"does it come with a charger?\". Answer it in the description.</li>\n          </ul>\n          <h2>Compatibility and fit</h2>\n          <p>For parts, accessories, consumables and refills, compatibility is the whole decision. A shopper asks \"will this filter fit my Model X200?\" and the agent needs a clear yes or no.</p>\n          <ul>\n            <li>List every compatible model by its full name and model number, not \"fits most models\".</li>\n            <li>List known incompatible models where confusion is likely, such as the Pro version of a device.</li>\n            <li>Put the list on the product page in text, not only in a separate compatibility tool that needs a form to be filled in.</li>\n          </ul>\n          <p>In structured data, schema.org has properties such as <code>isAccessoryOrSparePartFor</code> and <code>isConsumableFor</code> that point from your product to the product it's used with.</p>\n          <h2>Add the facts to structured data</h2>\n          <p>Product JSON-LD is where agents look first for price and stock. It can carry descriptive attributes too, so they don't have to be read out of your layout:</p>\n          <pre><code>{\n  \"@context\": \"https://schema.org\",\n  \"@type\": \"Product\",\n  \"name\": \"Ridgeline Packable Rain Jacket, Men's\",\n  \"sku\": \"NW-RJ-M-BLU\",\n  \"material\": \"100% recycled nylon\",\n  \"color\": \"Blue\",\n  \"size\": \"M\",\n  \"weight\": { \"@type\": \"QuantitativeValue\", \"value\": 410, \"unitCode\": \"GRM\" },\n  \"additionalProperty\": [\n    { \"@type\": \"PropertyValue\", \"name\": \"Waterproof rating\",\n      \"value\": 20000, \"unitText\": \"mm\" },\n    { \"@type\": \"PropertyValue\", \"name\": \"Packable\", \"value\": true }\n  ],\n  \"offers\": {\n    \"@type\": \"Offer\",\n    \"price\": \"189.00\",\n    \"priceCurrency\": \"USD\",\n    \"availability\": \"https://schema.org/InStock\"\n  }\n}</code></pre>\n          <p>Use <code>additionalProperty</code> for attributes schema.org has no dedicated property for. The values must match what the page says. For the full product markup, including variants, returns and shipping, see <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data that AI shopping agents can read</a>.</p>\n          <h2>Avoid vague copy</h2>\n          <p>Marketing copy still has a job: it persuades the person who reads it. The problem is when it replaces facts instead of sitting beside them. Watch for these:</p>\n          <ul>\n            <li><strong>Superlatives without evidence.</strong> \"Best-in-class\", \"industry-leading\", \"unbeatable\". An agent can't verify them, so it ignores them.</li>\n            <li><strong>Relative claims.</strong> \"30% lighter\" than what? Give the absolute number.</li>\n            <li><strong>Copy shared across a range.</strong> If ten products share the same paragraph, an agent can't tell them apart. Lead with what makes each one different.</li>\n            <li><strong>Manufacturer text pasted unchanged.</strong> It's often thin, and the same on every store selling that product. Add your own facts.</li>\n          </ul>\n          <p>A good pattern: one or two sentences of plain summary (what it is, who it's for, the one or two facts that matter most), then your brand copy, then the full specification table.</p>\n          <h2>What AgentScore checks</h2>\n          <p><a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> doesn't grade the wording of your descriptions. On a product page, it checks that schema.org Product data gives price, currency and availability (\"Product pages give price and stock in a form agents can read\"), and that prices are in the page HTML. To see how a real agent handles your descriptions, run a journey with <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a>, such as \"find a waterproof jacket under 500 grams\", and read the step-by-step replay. See <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">testing journeys with AI agents</a>.</p>\n          <h2>Where to start</h2>\n          <ol>\n            <li>Pick your best-selling category and list the five to ten attributes shoppers compare.</li>\n            <li>Audit ten product pages: is each attribute there, in text, with a unit?</li>\n            <li>Fill the gaps in your product information system or platform, not in each page by hand, so the data reaches every channel.</li>\n            <li>Add the specification table to the product page template and the same attributes to your Product JSON-LD.</li>\n            <li>Repeat for the next category.</li>\n          </ol>\n          <p>Shoppers also lean on prices, variants and reviews when comparing; see <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">machine-readable prices</a>, <a href=\"https://ghostagentlab.com/articles/variant-pickers-ai-agents/\">variant pickers AI agents can use</a> and <a href=\"https://ghostagentlab.com/articles/reviews-ratings-ai-agents/\">reviews and ratings AI agents can read</a>.</p>",
      "date_published": "2026-10-09T13:10:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/reviews-ratings-ai-agents/",
      "url": "https://ghostagentlab.com/articles/reviews-ratings-ai-agents/",
      "title": "Reviews and ratings AI agents can read",
      "summary": "How to make reviews and ratings visible to AI agents: reviews in the HTML, AggregateRating and Review markup that follows the rules, and honest moderation.",
      "content_html": "<p>\"Which of these has the best reviews?\" is one of the most natural questions to ask an AI shopping assistant. To answer it, the agent has to find your reviews, read them and trust them. Many stores load reviews in a way agents never see, or mark them up in ways that break the rules. Here's how to get it right.</p>\n          <h2>Why reviews matter to AI agents</h2>\n          <p>AI agents use reviews in two ways. The rating is a quick signal when they rank options: a 4.7 from several hundred reviews reads differently from a 4.7 from three. The review text answers specific questions the product description doesn't: \"does it run small?\", \"is it loud?\", \"how long did the battery last after a year?\".</p>\n          <p>If an agent can't see your reviews, it has nothing to weigh. It may fall back on reviews it finds elsewhere, or simply prefer a competitor whose ratings it can confirm.</p>\n          <h2>Put reviews in the HTML</h2>\n          <p>Most stores use a review app or service. Many of these load reviews with JavaScript after the page opens, from a separate server. A person sees them a moment later. An AI agent that reads the HTML without running JavaScript sees an empty space where the reviews should be, and many agents work that way.</p>\n          <ul>\n            <li><strong>Check the source.</strong> Open a product page, view the page source (not the browser's inspector, which shows the page after JavaScript), and search for a phrase from one of your reviews. If it isn't there, agents that skip JavaScript don't see it.</li>\n            <li><strong>Ask your provider about server-side output.</strong> Many review platforms offer a way to include the rating and the most recent or most helpful reviews in the page HTML, through a platform integration or an API. Turn it on if yours does.</li>\n            <li><strong>At minimum, render the summary.</strong> The average rating and review count, in text near the product name, plus a handful of reviews. \"Rated 4.6 out of 5 from 318 reviews\" is easy for any agent to read.</li>\n            <li><strong>Link paginated reviews.</strong> If reviews run to several pages, use ordinary links between them so agents can follow them.</li>\n          </ul>\n          <p>For the wider problem of content that only appears after JavaScript runs, see <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see JavaScript-only content</a>.</p>\n          <h2>Mark up ratings and reviews</h2>\n          <p>Schema.org gives agents the same facts in a standard form. On a product page, add <code>aggregateRating</code> for the summary and <code>review</code> for individual reviews, inside your Product JSON-LD:</p>\n          <pre><code>{\n  \"@context\": \"https://schema.org\",\n  \"@type\": \"Product\",\n  \"name\": \"Ethiopia Yirgacheffe Whole Bean Coffee, 340 g\",\n  \"aggregateRating\": {\n    \"@type\": \"AggregateRating\",\n    \"ratingValue\": \"4.6\",\n    \"bestRating\": \"5\",\n    \"reviewCount\": \"318\"\n  },\n  \"review\": [{\n    \"@type\": \"Review\",\n    \"author\": { \"@type\": \"Person\", \"name\": \"Dana R.\" },\n    \"datePublished\": \"2026-09-14\",\n    \"reviewRating\": { \"@type\": \"Rating\", \"ratingValue\": \"5\", \"bestRating\": \"5\" },\n    \"reviewBody\": \"Bright and floral as a pour-over. Arrived two days after roasting.\"\n  }],\n  \"offers\": {\n    \"@type\": \"Offer\",\n    \"price\": \"18.00\",\n    \"priceCurrency\": \"USD\",\n    \"availability\": \"https://schema.org/InStock\"\n  }\n}</code></pre>\n          <p>Use <code>reviewCount</code> for reviews with text, or <code>ratingCount</code> if you count star-only ratings too. Say what scale you use with <code>bestRating</code> if it isn't 1 to 5.</p>\n          <h2>The rules for review markup</h2>\n          <p>Search engines publish rules for review markup, and following them is the safest guide for AI agents too. The key ones from <a href=\"https://developers.google.com/search/docs/appearance/structured-data/review-snippet\">Google's review snippet guidelines</a>:</p>\n          <ul>\n            <li><strong>Only mark up reviews people can see.</strong> Every review and rating in the markup should be readable on the same page. Markup for reviews hidden somewhere else, or that don't exist, breaks the rules and can get your structured data ignored.</li>\n            <li><strong>Mark up reviews of a specific thing.</strong> A product, a recipe, a course, a book. Not a category page summing up many products.</li>\n            <li><strong>Don't add review markup about your own business on your own site.</strong> Google calls reviews that a business hosts about itself, on its own site, \"self-serving\", and doesn't show review stars for them on <code>LocalBusiness</code> and <code>Organization</code> pages. The safe reading is: put ratings on your products, and leave your company-level reputation to independent review sites.</li>\n            <li><strong>Keep the numbers in step.</strong> The rating and count in the markup should match what the page shows, and update when new reviews come in.</li>\n          </ul>\n          <p>Google's rules apply to its search results, not to every AI agent. But they describe what makes review data trustworthy, and agents that cross-check your markup against the page will notice the same problems.</p>\n          <h2>Honesty is the strategy</h2>\n          <p>AI agents read across many sources. If your site shows a perfect 5.0 and independent review sites say 3.2, the gap is itself a signal, and an agent may say so to the shopper. The only durable approach is reviews that reflect what customers actually think.</p>\n          <ul>\n            <li><strong>Publish the negative reviews.</strong> Hiding everything below four stars makes the remaining rating less believable, to people and to agents. In the US, the FTC's rule on fake reviews and testimonials covers suppressing negative reviews as well as faking positive ones. Check the rules where you sell.</li>\n            <li><strong>Label incentives.</strong> If a reviewer got the product free or was rewarded for reviewing, say so next to the review.</li>\n            <li><strong>Mark verified buyers.</strong> A \"verified purchase\" label, applied honestly, helps agents weigh reviews.</li>\n            <li><strong>Reply in public.</strong> A clear reply to a complaint (\"we've changed the packaging since March\") is information an agent can pass on.</li>\n            <li><strong>Don't pool unrelated products.</strong> Showing reviews of an old model, or of the whole range, on a new product's page misleads. Pool reviews only across true variants, like sizes and colors of the same item.</li>\n          </ul>\n          <h2>What AgentScore checks</h2>\n          <p><a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> doesn't have a dedicated reviews check. Its structured data checks look for schema.org JSON-LD on your home page, and for Product data with price, currency and availability on a product page, and flag markup that only appears after JavaScript runs. If your review app adds its markup with JavaScript, the same problem applies to your ratings. To see what an agent actually makes of your reviews, run a <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> journey such as \"find the best-rated coffee under $20 and add it to the cart\" and read the replay.</p>\n          <h2>A short checklist</h2>\n          <ol>\n            <li>View the source of a product page. Are the rating, review count and some review text in the HTML?</li>\n            <li>If not, turn on your review provider's server-side or SEO output, or ask your developer to render a summary.</li>\n            <li>Check the Product JSON-LD includes <code>aggregateRating</code> that matches the page.</li>\n            <li>Remove any rating markup about your business as a whole from your own site.</li>\n            <li>Review your moderation policy: are you publishing negative reviews and labeling incentives?</li>\n            <li>Validate with Google's Rich Results Test.</li>\n          </ol>\n          <p>Reviews work best alongside clear facts. See <a href=\"https://ghostagentlab.com/articles/product-descriptions-ai-agents/\">product descriptions AI agents can compare</a> and <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data that AI shopping agents can read</a>.</p>",
      "date_published": "2026-10-09T13:09:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/canonical-urls-ai-agents/",
      "url": "https://ghostagentlab.com/articles/canonical-urls-ai-agents/",
      "title": "Canonical URLs and duplicate pages for AI agents",
      "summary": "How to stop AI agents reading the same page at many URLs: rel=canonical, parameter and variant URLs, http, https and www redirects, and clean sitemaps.",
      "content_html": "<p>Most online stores show the same product at many different addresses: with tracking tags, with a color selected, under two categories, on http and https. People don't notice. AI agents do, because each address can look like a different page. Canonical URLs tell them which one is the real one.</p>\n          <h2>Why duplicate URLs confuse AI agents</h2>\n          <p>When an AI agent reads your site, or when an AI search service builds the index an assistant draws on, it meets your pages by URL. If the same jacket appears at five URLs, it may:</p>\n          <ul>\n            <li>treat them as five products, and compare your jacket with itself;</li>\n            <li>quote a price or stock level from an old copy it read weeks ago;</li>\n            <li>link a shopper to a version with a stale campaign tag, an odd filter or a variant they didn't ask for;</li>\n            <li>spend its limited time on your site reading copies instead of finding new products.</li>\n          </ul>\n          <p>The fix is not to make duplicates disappear, which is often impossible, but to say clearly which URL is the main one, and to make every other signal agree.</p>\n          <h2>Where duplicates come from</h2>\n          <table>\n            <thead><tr><th>Source</th><th>Example</th></tr></thead>\n            <tbody>\n              <tr><td>Tracking parameters</td><td><code>/products/rain-jacket?utm_source=newsletter</code></td></tr>\n              <tr><td>Sort and filter parameters</td><td><code>/jackets?sort=price&amp;color=blue</code></td></tr>\n              <tr><td>Variant parameters</td><td><code>/products/rain-jacket?variant=4417</code></td></tr>\n              <tr><td>Category paths</td><td><code>/mens/jackets/rain-jacket</code> and <code>/sale/rain-jacket</code></td></tr>\n              <tr><td>Protocol and host</td><td><code>http://</code> and <code>https://</code>, with and without <code>www.</code></td></tr>\n              <tr><td>Trailing slashes and case</td><td><code>/Rain-Jacket/</code> and <code>/rain-jacket</code></td></tr>\n            </tbody>\n          </table>\n          <h2>Set a canonical URL on every page</h2>\n          <p>A canonical link is one line in the page's <code>&lt;head&gt;</code> that names the preferred URL for the content:</p>\n          <pre><code>&lt;link rel=\"canonical\" href=\"https://northwind.example/products/rain-jacket\"&gt;</code></pre>\n          <p>Put it on every indexable page, including on the canonical page itself (a self-referencing canonical). Rules that keep it reliable:</p>\n          <ul>\n            <li><strong>Use the full, absolute URL.</strong> Include <code>https://</code> and the host, not just the path.</li>\n            <li><strong>Put it in the HTML.</strong> A canonical tag added by JavaScript may never be seen by agents and crawlers that read the raw page.</li>\n            <li><strong>One per page.</strong> Two different canonical tags on a page is a common theme or app conflict. Search engines may ignore both.</li>\n            <li><strong>Point to a page that works.</strong> The canonical should return 200, not redirect, 404 or carry a <code>noindex</code>.</li>\n            <li><strong>Don't canonicalize everything to the home page.</strong> Each distinct page should point at itself, or at its true duplicate.</li>\n          </ul>\n          <p>Search engines treat <code>rel=\"canonical\"</code> as a strong hint rather than a command, and the same is likely true of AI services that build on them. That's why the other signals below need to agree with it.</p>\n          <h2>Parameter URLs</h2>\n          <p>Tracking parameters like <code>utm_source</code> and click IDs never change the content, so the page should always name the clean URL as canonical. Sort orders and session IDs are the same.</p>\n          <p>Filters are a judgment call. A filtered listing such as \"blue jackets\" may be worth keeping as its own page if people search for it. If so, give it a clean path (<code>/jackets/blue</code>), its own title and a self-referencing canonical. Most combinations of filters, though, should point back to the unfiltered category. For more on filters agents can use, see <a href=\"https://ghostagentlab.com/articles/site-search-filters-ai-agents/\">site search and filters for AI agents</a>.</p>\n          <h2>Product variants</h2>\n          <p>Variants are where stores most often go wrong. There are two reasonable patterns:</p>\n          <ul>\n            <li><strong>One page for the product.</strong> Every variant URL (<code>?variant=4417</code>, <code>?color=blue</code>) has a canonical pointing to the main product URL. This suits products where variants differ only in size or color and share a description. Make sure the main page lists every variant's price and availability, ideally in <code>ProductGroup</code> markup, so the agent doesn't lose them.</li>\n            <li><strong>A page per variant.</strong> Each variant has its own URL with a self-referencing canonical, its own title (\"Rain Jacket, Men's, Blue\") and its own price. This suits variants that are really different products, with different specs or prices.</li>\n          </ul>\n          <p>What doesn't work is a mix: variant URLs that canonicalize to the main page, while the main page only shows the default variant's price and stock. An agent is then told the blue jacket is the same page as the red one, but can't find the blue one's details. See <a href=\"https://ghostagentlab.com/articles/variant-pickers-ai-agents/\">size, color and variant pickers AI agents can use</a>.</p>\n          <h2>http, https and www: redirect, don't just canonicalize</h2>\n          <p>For duplicates that are purely technical, a canonical tag isn't enough. Send a permanent (301 or 308) redirect so that every visitor, person or agent, lands on one version:</p>\n          <ul>\n            <li><code>http://</code> redirects to <code>https://</code>.</li>\n            <li>Pick <code>www.northwind.example</code> or <code>northwind.example</code>, and redirect the other.</li>\n            <li>Pick trailing slash or no trailing slash, and redirect the other.</li>\n            <li>Redirect old product URLs to their replacements after a migration, not to the home page.</li>\n          </ul>\n          <p>Redirect in a single hop. Chains (http to https, then to www, then to a trailing slash) slow every request down, and some clients stop following after a few hops. Avoid redirects that depend on JavaScript or a meta refresh; agents that don't run scripts get stuck on the first page. And make sure region or language redirects don't trap agents, as covered in <a href=\"https://ghostagentlab.com/articles/international-sites-ai-agents/\">international sites for AI agents</a>.</p>\n          <h2>Make every signal agree</h2>\n          <p>Agents and crawlers look at more than the canonical tag. Each of these should use the same, canonical URL:</p>\n          <ul>\n            <li><strong>Your XML sitemap.</strong> List only canonical URLs that return 200. No parameter URLs, no redirecting URLs, no http versions. See <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">XML sitemaps for AI agents</a>.</li>\n            <li><strong>Internal links.</strong> Navigation, product grids and breadcrumbs should link to the canonical URL directly, not to a version that redirects or carries parameters.</li>\n            <li><strong>Structured data.</strong> The <code>url</code> in your Product and Offer JSON-LD should match the canonical.</li>\n            <li><strong><code>og:url</code> and feeds.</strong> Social tags and product feeds should use the same address.</li>\n            <li><strong>hreflang tags</strong>, if you have them, should point to canonical URLs of each language version.</li>\n          </ul>\n          <h2>What AgentScore checks</h2>\n          <p><a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> doesn't have a dedicated canonical check. The \"Valid sitemap\" check confirms you have a sitemap it can read (found through robots.txt, <code>/sitemap.xml</code> or <code>/sitemap_index.xml</code>) and counts its entries; it doesn't judge whether each URL is canonical. AgentScore also uses your sitemap and home page links to find a product page to test, so a sitemap full of redirecting or parameter URLs makes that harder for any agent.</p>\n          <h2>A quick audit</h2>\n          <ol>\n            <li>Open a product page with <code>?utm_source=test</code> on the end. View the source and check the canonical points to the clean URL.</li>\n            <li>Do the same for a variant URL and a filtered category. Is the result the pattern you intended?</li>\n            <li>Type <code>http://</code> and the non-preferred host into a browser. Do they each redirect, in one hop, to the right page?</li>\n            <li>Open your sitemap and spot-check ten URLs: do they all return 200, and match their own canonical?</li>\n            <li>Check your main navigation links for parameters or redirects.</li>\n          </ol>\n          <p>Canonical URLs are housekeeping, but they decide which version of your page agents read and repeat. For the bigger picture, see the <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness guide</a>.</p>",
      "date_published": "2026-10-09T13:08:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/international-sites-ai-agents/",
      "url": "https://ghostagentlab.com/articles/international-sites-ai-agents/",
      "title": "Languages, currencies and regions: international sites for AI agents",
      "summary": "How to make international stores work for AI agents: lang and hreflang, unambiguous currencies, region redirects that don't trap agents, and shipping countries.",
      "content_html": "<p>If you sell in more than one country, an AI agent needs to know which version of your site it's reading: what language it's in, which currency the prices are in, and whether you'll deliver to the shopper's address. Get this wrong and an agent may quote a euro price to someone in Ohio, or never get past your country picker at all.</p>\n          <h2>What can go wrong</h2>\n          <p>International sites are built for people, who can see a flag in the corner, read a price symbol and click \"change country\". AI agents work from the page's code and text, often from a server in another country, and usually without the browser settings that a person's visit carries. Common failures:</p>\n          <ul>\n            <li>The agent is redirected to the wrong country store based on where its server is, and can't get back.</li>\n            <li>A price shows \"$\" with no sign of whether it's US, Canadian or Australian dollars.</li>\n            <li>Three language versions of the same page look like three unrelated pages, or like duplicates.</li>\n            <li>The agent tells a shopper a product is available, when you don't ship to their country.</li>\n          </ul>\n          <h2>Declare the language with the lang attribute</h2>\n          <p>Every page should say what language it's in, on its <code>&lt;html&gt;</code> tag:</p>\n          <pre><code>&lt;html lang=\"en-GB\"&gt;</code></pre>\n          <p>Use a language code, optionally with a region: <code>en</code>, <code>en-GB</code>, <code>fr-CA</code>, <code>de-CH</code>. It tells agents, screen readers and translation tools how to read the page. Make sure it changes with the language version. A French page that still says <code>lang=\"en\"</code>, because the theme hard-codes it, is surprisingly common. The <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> check \"Clear page title, description, and headings\" flags a missing <code>lang</code> attribute on your home page; see <a href=\"https://ghostagentlab.com/articles/titles-descriptions-headings-ai/\">titles, descriptions and headings</a>.</p>\n          <h2>Connect language and country versions with hreflang</h2>\n          <p><code>hreflang</code> links tell agents and search engines that several pages are versions of the same content for different languages or regions, and which one is for whom. Each version lists all the others, and itself:</p>\n          <pre><code>&lt;link rel=\"alternate\" hreflang=\"en-us\" href=\"https://northwind.example/us/products/rain-jacket\"&gt;\n&lt;link rel=\"alternate\" hreflang=\"en-gb\" href=\"https://northwind.example/uk/products/rain-jacket\"&gt;\n&lt;link rel=\"alternate\" hreflang=\"fr-fr\" href=\"https://northwind.example/fr/produits/veste-pluie\"&gt;\n&lt;link rel=\"alternate\" hreflang=\"x-default\" href=\"https://northwind.example/products/rain-jacket\"&gt;</code></pre>\n          <ul>\n            <li><strong>Make it two-way.</strong> If the US page lists the UK page, the UK page must list the US page. One-way links are usually ignored.</li>\n            <li><strong>Use canonical URLs.</strong> Each <code>href</code> should be the canonical URL of that version, returning 200, not a redirect. See <a href=\"https://ghostagentlab.com/articles/canonical-urls-ai-agents/\">canonical URLs and duplicate pages</a>.</li>\n            <li><strong>Don't canonicalize across languages.</strong> The French page's canonical should point to the French page, not the English one.</li>\n            <li><strong>Add <code>x-default</code>.</strong> It names the page for everyone else, typically a global version or a country picker that works without redirects.</li>\n            <li><strong>Use correct codes.</strong> Language first, then optional region: <code>en-gb</code>, not <code>uk</code> or <code>gb-en</code>.</li>\n          </ul>\n          <p>You can put hreflang in the page <code>&lt;head&gt;</code>, in HTTP headers, or in your XML sitemap. Pick one method and keep it consistent.</p>\n          <h2>Region redirects that trap agents</h2>\n          <p>Many stores redirect visitors automatically based on their IP address or browser language. For people, it saves a click. For AI agents it can be a dead end. Agents often run from cloud data centers, so the location you detect is the server's, not the shopper's. An agent helping someone in Berlin may be sent to your US store, and if every attempt to open <code>/de/</code> bounces back to <code>/us/</code>, it can never read the German prices.</p>\n          <ul>\n            <li><strong>Suggest, don't force.</strong> Show a banner (\"Looks like you're in Germany. Go to the German store?\") instead of redirecting.</li>\n            <li><strong>Never redirect away from an explicit country URL.</strong> If a request asks for <code>/de/products/rain-jacket</code>, serve it.</li>\n            <li><strong>Let the choice stick in the URL.</strong> Country and language in the path or subdomain work for agents. A choice stored only in a cookie doesn't, because many agents don't keep cookies between requests.</li>\n            <li><strong>Don't block whole regions without a reason.</strong> Blocking traffic from countries you don't sell to can also block the data centers AI agents run from. See <a href=\"https://ghostagentlab.com/articles/geo-blocking-ai-agents/\">geo-blocking, VPN blocks and AI agents</a>.</li>\n          </ul>\n          <h2>Make the currency unambiguous</h2>\n          <p>A symbol alone is often not enough. \"$\" is used by more than a dozen currencies, and \"kr\" by several. Show the currency code where there's any doubt (<code>US$189</code> or <code>189.00 USD</code>), and always put it in your structured data:</p>\n          <pre><code>\"offers\": {\n  \"@type\": \"Offer\",\n  \"price\": \"175.00\",\n  \"priceCurrency\": \"GBP\",\n  \"availability\": \"https://schema.org/InStock\"\n}</code></pre>\n          <p><code>priceCurrency</code> takes a three-letter ISO 4217 code: <code>USD</code>, <code>GBP</code>, <code>EUR</code>, <code>CAD</code>. Each country version should show its own price and currency in both the page and the markup. Say whether prices include sales tax or VAT, since that differs by country. More in <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">machine-readable prices</a>.</p>\n          <div>\n            <p><strong>Watch out for converted prices.</strong> If prices are converted on the fly with JavaScript, an agent that doesn't run JavaScript sees your base currency, while the shopper would have paid something else. Render each store's price in the HTML.</p>\n          </div>\n          <h2>Say where you ship</h2>\n          <p>\"Do they deliver to Canada?\" is a question an agent should be able to answer before sending someone to your checkout. Make it easy:</p>\n          <ul>\n            <li><strong>A shipping page in plain text</strong> listing the countries you ship to, with costs and delivery times for each. Not only a dropdown at checkout.</li>\n            <li><strong>Product-level restrictions on the product page.</strong> If an item can't ship to some countries, say so on that page.</li>\n            <li><strong>Shipping details in structured data.</strong> Schema.org's <code>shippingDetails</code> on an Offer can describe where you deliver, at what cost, and how long it takes:</li>\n          </ul>\n          <pre><code>\"shippingDetails\": {\n  \"@type\": \"OfferShippingDetails\",\n  \"shippingDestination\": { \"@type\": \"DefinedRegion\", \"addressCountry\": [\"GB\", \"IE\"] },\n  \"shippingRate\": { \"@type\": \"MonetaryAmount\", \"value\": \"4.95\", \"currency\": \"GBP\" },\n  \"deliveryTime\": {\n    \"@type\": \"ShippingDeliveryTime\",\n    \"handlingTime\": { \"@type\": \"QuantitativeValue\", \"minValue\": 0, \"maxValue\": 1, \"unitCode\": \"DAY\" },\n    \"transitTime\": { \"@type\": \"QuantitativeValue\", \"minValue\": 2, \"maxValue\": 4, \"unitCode\": \"DAY\" }\n  }\n}</code></pre>\n          <p>Returns can differ by country too; give each <code>MerchantReturnPolicy</code> its <code>applicableCountry</code>. See <a href=\"https://ghostagentlab.com/articles/organization-structured-data/\">Organization and LocalBusiness data</a> and <a href=\"https://ghostagentlab.com/articles/policy-pages-ai-assistants/\">policy pages AI assistants can quote</a>.</p>\n          <h2>A checklist for each country version</h2>\n          <ol>\n            <li>The <code>&lt;html lang&gt;</code> attribute matches the page's language.</li>\n            <li>hreflang links connect every version both ways, with an <code>x-default</code>.</li>\n            <li>Each version's canonical points to itself.</li>\n            <li>Opening a country URL directly serves that country's page, with no forced redirect.</li>\n            <li>Prices are in the HTML, with an unambiguous currency, and <code>priceCurrency</code> matches.</li>\n            <li>The countries you ship to are listed in text, and in <code>shippingDetails</code> where you can.</li>\n          </ol>\n          <p>If your country stores run on separate domains, such as <code>northwind.example</code> and <code>northwind-uk.example</code>, scan each one with <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> to compare them. To check a real journey in another market, such as \"find a rain jacket and get to checkout on the UK store\", run it with <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> and read the replay.</p>",
      "date_published": "2026-10-09T13:07:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/variant-pickers-ai-agents/",
      "url": "https://ghostagentlab.com/articles/variant-pickers-ai-agents/",
      "title": "Size, color and variant pickers AI agents can use",
      "summary": "Build size and color pickers AI agents can use: named options, a readable selected state, clear sold-out options, variant URLs and prices that update.",
      "content_html": "<p>Ask an AI agent to buy \"the navy jacket in a medium\" and the hard part isn't finding the jacket. It's choosing navy and medium. Size, color and other variant pickers are where many product pages lose agents: the swatch has no name, the selected size looks different but reads the same, and the sold-out option can still be clicked. This guide covers how to build variant pickers AI agents can use without guessing.</p>\n          <h2>Why variant pickers trip agents up</h2>\n          <p>A person sees a row of colored circles, notices a thicker border around one, and sees a crossed-out \"XL\". A browser agent mostly reads the page's structure: each control's role, its name and its state. If the swatch is a <code>&lt;div&gt;</code> with a background color, it has no role and no name. If \"selected\" is shown only with a border, there's no state to read. The agent may click the right circle by chance, click the wrong one, or not realize there's a choice to make at all.</p>\n          <p>The cost is direct. An agent that can't choose a size either gives up and tells the shopper it couldn't complete the task, or adds the wrong item to the cart. Both are lost or returned orders. And the same problems affect shoppers who use a keyboard or a screen reader, so fixing them helps people too.</p>\n          <p>Agents need four things from a picker: what each option is called, which one is chosen, which ones can't be bought, and confirmation that the page changed after they chose.</p>\n          <h2>1. Make each option a real control with a name</h2>\n          <p>The simplest, most reliable picker is a group of native radio buttons. Only one option can be chosen, the browser handles the keyboard, and every agent knows what a radio button is. You can style the label to look like a swatch or a size box.</p>\n          <pre><code>&lt;fieldset&gt;\n  &lt;legend&gt;Color&lt;/legend&gt;\n  &lt;label class=\"swatch\"&gt;\n    &lt;input type=\"radio\" name=\"color\" value=\"navy\" checked&gt;\n    &lt;span class=\"dot\" style=\"background:#1f2a44\" aria-hidden=\"true\"&gt;&lt;/span&gt;\n    Navy\n  &lt;/label&gt;\n  &lt;label class=\"swatch\"&gt;\n    &lt;input type=\"radio\" name=\"color\" value=\"olive\"&gt;\n    &lt;span class=\"dot\" style=\"background:#5b5e3a\" aria-hidden=\"true\"&gt;&lt;/span&gt;\n    Olive\n  &lt;/label&gt;\n&lt;/fieldset&gt;</code></pre>\n          <ul>\n            <li><strong>Name every option.</strong> The color name, such as \"Navy\", should be the control's name, as visible text or at least an <code>aria-label</code>. A hex code or an image file name isn't a name.</li>\n            <li><strong>Name the group.</strong> A <code>&lt;fieldset&gt;</code> with a <code>&lt;legend&gt;</code> (\"Color\", \"Size\") tells the agent what the choice is about.</li>\n            <li><strong>Use the words shoppers use.</strong> \"Medium\" or \"M\" is fine. An internal code like \"SZ-03\" isn't. If sizes need context (\"UK 8 / US 4\"), put it in the name.</li>\n            <li><strong>Keep swatch images as decoration.</strong> If a swatch is a photo of the fabric, give the control the color name and the image empty <code>alt=\"\"</code>, so the name isn't read twice.</li>\n          </ul>\n          <p>If your design uses buttons rather than radios, that works too, as long as they are real <code>&lt;button&gt;</code> elements with names. AgentScore's checks that buttons and links have names, and that clickable things are real buttons and links, look for exactly these problems, though they run on your home page. A product page needs the same care.</p>\n          <h2>2. Show which option is selected, in the code</h2>\n          <p>Selection is the detail most often shown only visually. Radio buttons report it for free through <code>checked</code>. Custom buttons have to say it themselves:</p>\n          <pre><code>&lt;div role=\"group\" aria-labelledby=\"size-label\"&gt;\n  &lt;span id=\"size-label\"&gt;Size&lt;/span&gt;\n  &lt;button type=\"button\" aria-pressed=\"false\"&gt;S&lt;/button&gt;\n  &lt;button type=\"button\" aria-pressed=\"true\"&gt;M&lt;/button&gt;\n  &lt;button type=\"button\" aria-pressed=\"false\"&gt;L&lt;/button&gt;\n&lt;/div&gt;</code></pre>\n          <ul>\n            <li>Use <code>aria-pressed=\"true\"</code> on the chosen button and <code>\"false\"</code> on the others, and update them when the choice changes.</li>\n            <li>Alternatively, use <code>role=\"radio\"</code> with <code>aria-checked</code> inside a <code>role=\"radiogroup\"</code>. It's more work to get the keyboard right, which is why native radios are usually the better choice.</li>\n            <li>Repeat the current choice in text near the picker (\"Color: Navy\"). It gives agents and people a plain confirmation.</li>\n          </ul>\n          <h2>3. Mark out-of-stock options clearly</h2>\n          <p>Hiding a sold-out size makes the shopper wonder whether it ever existed. Showing it with only a strike-through or faded color leaves an agent free to choose it, then fail at \"Add to cart\". Say it in words and in the control's state.</p>\n          <table>\n            <thead><tr><th>Approach</th><th>What an agent sees</th></tr></thead>\n            <tbody>\n              <tr><td>Option faded or struck through with CSS only</td><td>A normal option it can pick</td></tr>\n              <tr><td>Option removed from the page</td><td>No sign the size exists</td></tr>\n              <tr><td>Native radio with <code>disabled</code> and \"XL, sold out\" in the label</td><td>An option that can't be chosen, and why</td></tr>\n              <tr><td>Button with <code>aria-disabled=\"true\"</code> and \"XL, sold out\" in its name</td><td>The same, and it stays reachable by keyboard</td></tr>\n            </tbody>\n          </table>\n          <p>If you offer back-in-stock alerts, make that a separate, named button (\"Email me when XL is back\"), so an agent can tell it apart from buying.</p>\n          <h2>4. Give each variant its own URL</h2>\n          <p>Many AI assistants fetch pages rather than click through them. They can't press a swatch, but they can open a link. If each variant has its own address, an assistant can point a shopper straight to \"the navy jacket in medium\", and a browser agent can recover if a click doesn't take.</p>\n          <pre><code>https://northwind.example/products/field-jacket?color=navy&amp;size=m\nhttps://northwind.example/products/field-jacket?variant=40318</code></pre>\n          <ul>\n            <li>Update the address when a shopper picks an option, so the current view always has a link.</li>\n            <li>Make the URL load with that variant already selected, its price and stock shown, when opened fresh.</li>\n            <li>Point the canonical URL at the main product page unless variants are really separate products. See <a href=\"https://ghostagentlab.com/articles/canonical-urls-ai-agents/\">canonical URLs and duplicate pages</a>.</li>\n          </ul>\n          <h2>5. Update price, stock and photos where agents can see it</h2>\n          <p>Choosing a larger size or a different material often changes the price. If the new price only appears in an image, or changes quietly in a corner of the page, an agent may report the old one to the shopper.</p>\n          <ul>\n            <li>Show price and availability as text, and update them when the variant changes.</li>\n            <li>Wrap the price and stock line in a live region (<code>aria-live=\"polite\"</code>), so the change is announced to assistive tech and is easier for agents reading the page structure to notice.</li>\n            <li>Keep \"Add to cart\" a real button with the same name throughout. If it's disabled until a size is chosen, say so next to it: \"Choose a size\". AgentScore's add-to-cart check accepts a button that's disabled until options are chosen, because agents can handle that, as long as they can tell what's missing.</li>\n          </ul>\n          <h2>6. Describe variants in your structured data</h2>\n          <p>Agents that read your page HTML rather than click it rely on structured data for the full list of options. schema.org has a <a href=\"https://schema.org/ProductGroup\">ProductGroup</a> type for this: one group, with <code>variesBy</code> naming the dimensions and <code>hasVariant</code> listing each variant as a <code>Product</code> with its own offer. Google documents the pattern in its <a href=\"https://developers.google.com/search/docs/appearance/structured-data/product-variants\">product variant structured data guide</a>.</p>\n          <pre><code>{\n  \"@context\": \"https://schema.org\",\n  \"@type\": \"ProductGroup\",\n  \"name\": \"Field jacket\",\n  \"productGroupID\": \"FJ-100\",\n  \"variesBy\": [\"https://schema.org/color\", \"https://schema.org/size\"],\n  \"hasVariant\": [{\n    \"@type\": \"Product\",\n    \"sku\": \"FJ-100-NAV-M\",\n    \"color\": \"Navy\",\n    \"size\": \"M\",\n    \"url\": \"https://northwind.example/products/field-jacket?variant=40318\",\n    \"offers\": {\n      \"@type\": \"Offer\",\n      \"price\": \"189.00\",\n      \"priceCurrency\": \"USD\",\n      \"availability\": \"https://schema.org/InStock\"\n    }\n  }]\n}</code></pre>\n          <p>AgentScore's \"Product pages give price and stock in a form agents can read\" check reads offers inside <code>ProductGroup</code> and <code>hasVariant</code>, as well as on a plain <code>Product</code>. For more on the fields that matter, see <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data AI shopping agents can read</a> and <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">prices AI agents can read</a>.</p>\n          <h2>How to test your pickers</h2>\n          <ol>\n            <li><strong>Use only the keyboard.</strong> Tab to the color and size pickers, choose an option, and add to cart. If you can't, many agents can't either.</li>\n            <li><strong>Inspect the accessibility tree.</strong> In Chrome's developer tools, the Accessibility pane should show each option as a radio or button with a name like \"Navy\", and its checked or pressed state.</li>\n            <li><strong>Pick a sold-out size.</strong> The page should tell you, in words, before you reach the cart.</li>\n            <li><strong>Copy the URL after choosing.</strong> Open it in a private window. The same variant should be selected.</li>\n            <li><strong>Run a journey.</strong> AgentScore's free scan doesn't choose options on your product pages, so the best test is an agent trying it. A <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> can run \"find the field jacket, choose navy in medium, add it to the cart\" on a schedule and replay each step when it gets stuck. See <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">testing key journeys with AI agents</a>.</li>\n          </ol>\n          <div>\n            <p><strong>Where to start.</strong> If your theme uses <code>&lt;div&gt;</code> swatches, switching to labelled radio buttons styled the same way usually fixes naming, selection state and keyboard use in one change. Ask your developer or theme provider for that first.</p>\n          </div>\n          <p>For the rest of the page's controls, see <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>. For what happens after \"Add to cart\", see <a href=\"https://ghostagentlab.com/articles/cart-ai-agents/\">carts AI agents can manage</a>.</p>",
      "date_published": "2026-10-09T13:06:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Navigability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/infinite-scroll-pagination-ai-agents/",
      "url": "https://ghostagentlab.com/articles/infinite-scroll-pagination-ai-agents/",
      "title": "Infinite scroll, “load more” and pagination for AI agents",
      "summary": "Infinite scroll can hide most of your catalog from AI agents. Back it with real paginated URLs, a “load more” link and a sitemap.",
      "content_html": "<p>Infinite scroll feels effortless to a shopper with a thumb. To many AI agents, it's a wall. Products that only load as someone scrolls may never be seen by an assistant that fetches the page, and a browser agent can waste its time scrolling and waiting. The fix isn't to drop infinite scroll or \"load more\". It's to put real, linked pages underneath them.</p>\n          <h2>What agents see on a long list</h2>\n          <p>Different agents meet a category page in different ways:</p>\n          <ul>\n            <li><strong>Assistants that fetch pages</strong>, such as ChatGPT or Perplexity looking something up for a person, usually read the HTML your server sends. They don't scroll, and often don't run JavaScript. They see the first batch of products and nothing more.</li>\n            <li><strong>AI search crawlers</strong> also read HTML and follow links. If page two has no link, it doesn't exist for them.</li>\n            <li><strong>Browser agents</strong> can scroll and click, but each scroll is a step that takes time. They can't always tell whether more products are coming or the list has ended.</li>\n          </ul>\n          <p>So a category of 300 products that shows 24 at a time, with more loading on scroll, can look like a category of 24. If the product a shopper asked for is number 61, the assistant may tell them you don't sell it.</p>\n          <h2>Give every page of results its own URL</h2>\n          <p>The foundation is the same for infinite scroll, \"load more\" and numbered pages: each chunk of results should exist at its own address, and that address should load those results when opened fresh.</p>\n          <pre><code>https://northwind.example/collections/coffee\nhttps://northwind.example/collections/coffee?page=2\nhttps://northwind.example/collections/coffee?page=3</code></pre>\n          <ul>\n            <li><strong>Use query parameters or paths, not fragments.</strong> <code>?page=2</code> and <code>/page/2</code> are separate URLs. <code>#page=2</code> is not: servers never see what follows the <code>#</code>, so it can't return different results.</li>\n            <li><strong>Render the results in the HTML.</strong> <code>?page=2</code> should arrive with its products already in it, not an empty shell that fills in later. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see your JavaScript-only content</a>.</li>\n            <li><strong>Keep the order stable.</strong> Page two should show the same products whether you arrive from page one or open it directly. Random or personalized ordering makes pages unreliable to follow.</li>\n            <li><strong>Carry filters and sort along.</strong> <code>?roast=dark&amp;sort=price-asc&amp;page=2</code> should keep both. Our guide to <a href=\"https://ghostagentlab.com/articles/site-search-filters-ai-agents/\">site search and filters AI agents can use</a> covers filter URLs in detail.</li>\n          </ul>\n          <p>Google's guidance on <a href=\"https://developers.google.com/search/docs/specialty/ecommerce/pagination-and-incremental-page-loading\">pagination and incremental page loading</a> takes the same approach, and it works for AI agents for the same reasons.</p>\n          <h2>Link the pages together with real links</h2>\n          <p>Unique URLs only help if agents can find them. Each page should link to the next one, and ideally to the previous one and a few page numbers, with ordinary <code>&lt;a href&gt;</code> links in the HTML.</p>\n          <pre><code>&lt;nav aria-label=\"Pagination\"&gt;\n  &lt;a href=\"/collections/coffee?page=1\"&gt;Previous page&lt;/a&gt;\n  &lt;a href=\"/collections/coffee?page=1\"&gt;1&lt;/a&gt;\n  &lt;a href=\"/collections/coffee?page=2\" aria-current=\"page\"&gt;2&lt;/a&gt;\n  &lt;a href=\"/collections/coffee?page=3\"&gt;3&lt;/a&gt;\n  &lt;a href=\"/collections/coffee?page=3\"&gt;Next page&lt;/a&gt;\n&lt;/nav&gt;</code></pre>\n          <ul>\n            <li>Wrap the links in <code>&lt;nav&gt;</code> with a name like \"Pagination\", so agents can tell it apart from the main menu.</li>\n            <li>Name arrow links in words (\"Next page\"), not just \"›\".</li>\n            <li>Mark the current page with <code>aria-current=\"page\"</code>.</li>\n            <li>Say how many products there are in text, such as \"Showing 25–48 of 312\". It tells an agent there's more to see.</li>\n          </ul>\n          <h2>Make \"load more\" a button, backed by a link</h2>\n          <p>\"Load more\" is a good middle ground for shoppers, and it can work well for agents if it's built as two things at once: a real button for people and browser agents, and a real link for everyone else.</p>\n          <pre><code>&lt;a class=\"load-more\" href=\"/collections/coffee?page=3\"&gt;\n  Load more products (264 remaining)\n&lt;/a&gt;</code></pre>\n          <p>Start with a plain link to the next page. Then let JavaScript take over: intercept the click, fetch the next page, append the products, and update the address with <code>history.pushState</code> so the URL matches what's on screen. If the script doesn't run, or the agent doesn't run scripts, the link still goes to page three. If you prefer a <code>&lt;button&gt;</code>, keep the numbered page links alongside it so there's still a route.</p>\n          <ul>\n            <li>Give the button a specific name: \"Load more products\", not \"More\" or an arrow icon.</li>\n            <li>Announce new results with a live region (\"24 more products loaded\") so the change is easier for agents and assistive tech to notice.</li>\n            <li>Move keyboard focus to the first new product, or leave it on the button, but don't send it back to the top of the page.</li>\n          </ul>\n          <h2>Infinite scroll: keep it, but add pages underneath</h2>\n          <p>If you want true infinite scroll, treat it as an enhancement over paginated pages. As each new batch loads, update the URL to the page number the shopper has reached. Keep the paginated links in the HTML, even if they're visually tucked away at the bottom. And make sure the footer is reachable: if the page grows every time someone nears the bottom, neither people nor agents can reach your contact, shipping and returns links.</p>\n          <table>\n            <thead><tr><th>Pattern</th><th>Fetching assistants and crawlers</th><th>Browser agents</th></tr></thead>\n            <tbody>\n              <tr><td>Infinite scroll only</td><td>See the first batch only</td><td>Must scroll and wait, can't tell when it ends</td></tr>\n              <tr><td>\"Load more\" button with no URL</td><td>See the first batch only</td><td>Can click, if the button has a name</td></tr>\n              <tr><td>\"Load more\" link to <code>?page=2</code></td><td>Can follow every page</td><td>Can click or follow</td></tr>\n              <tr><td>Numbered page links</td><td>Can follow every page</td><td>Can jump straight to a page</td></tr>\n            </tbody>\n          </table>\n          <h2>What about rel=\"next\" and rel=\"prev\"?</h2>\n          <p>For years, <code>&lt;link rel=\"next\"&gt;</code> and <code>rel=\"prev\"</code> in the page head were the standard way to tell search engines that pages belonged in a series. Google said in 2019 that it no longer uses them for indexing. Other search engines and tools may still read them, and they're part of HTML, so they do no harm. But don't rely on them: an agent following links needs the visible <code>&lt;a href&gt;</code> links in the page.</p>\n          <p>Two related head tags matter more:</p>\n          <ul>\n            <li><strong>Canonical.</strong> Give each paginated page its own canonical URL. Pointing <code>?page=2</code> at page one tells crawlers to ignore the products on page two. See <a href=\"https://ghostagentlab.com/articles/canonical-urls-ai-agents/\">canonical URLs and duplicate pages</a>.</li>\n            <li><strong>Robots meta.</strong> Don't mark paginated pages <code>noindex</code> by default. Pages that stay out of the index can, over time, get less attention from crawlers, and so can the products linked from them.</li>\n          </ul>\n          <h2>Back it up with a sitemap</h2>\n          <p>Even with perfect pagination, some products sit deep in a long list. Your XML sitemap should list every product page directly, so crawlers don't depend on paging through categories to find them. AgentScore's \"Valid sitemap\" check confirms you have one. See <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">XML sitemaps for AI agents</a>. The sitemap is the safety net. The links are still the main route.</p>\n          <h2>How to check your category pages</h2>\n          <ol>\n            <li><strong>View the source</strong> of a category page and search for a product name from further down the list. Then look for a link to page two.</li>\n            <li><strong>Open <code>?page=2</code> directly</strong> in a private window. It should show the second set of products, not the first and not an empty page.</li>\n            <li><strong>Turn off JavaScript</strong> and try to reach the last page of a category.</li>\n            <li><strong>Check the footer</strong> is reachable on a category page with infinite scroll.</li>\n            <li><strong>Run AgentScore.</strong> It has no pagination check, but \"Content loads without JavaScript\" and \"Key pages are linked from the home page\" show whether HTML-only agents can see your content and reach your products.</li>\n          </ol>\n          <div>\n            <p><strong>Test it as a shopper would.</strong> Ask a <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> to \"find the cheapest dark roast in the coffee category\" on a category of more than one page. The step-by-step replay shows whether it reached the later pages or settled for what loaded first.</p>\n          </div>\n          <p>For the bigger picture of how agents move around a site, see <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">What is agent readiness?</a></p>",
      "date_published": "2026-10-09T13:05:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Navigability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/error-messages-ai-agents/",
      "url": "https://ghostagentlab.com/articles/error-messages-ai-agents/",
      "title": "Error messages AI agents can understand and fix",
      "summary": "Help AI agents recover from mistakes: error text tied to each field, live regions, useful 404 pages and HTTP status codes that match what the page says.",
      "content_html": "<p>Every journey hits an error eventually: a mistyped postcode, a promo code that's expired, a product that's gone. A person reads the red text, fixes it and carries on. An AI agent can only do the same if the error is written down, attached to the right field and reported honestly by your server. Otherwise it tries again, gives up, or tells the shopper something went wrong without saying what.</p>\n          <h2>Why errors matter more for agents</h2>\n          <p>An agent working for someone wants to finish the task. When a form doesn't go through, it needs to answer three questions: did something go wrong, which field caused it, and what should it enter instead? People answer those from visual cues, such as a red border, a shake, an icon or a message that flashes up and disappears. Agents mostly read the page's text and structure, and may not catch any of those cues.</p>\n          <p>A clear error turns a failed attempt into a quick correction. An unclear one ends the journey, and you lose the order or sign-up without anyone telling you. The same fixes help people who use screen readers, and anyone filling in a form on a small screen.</p>\n          <h2>Write errors that say what to do</h2>\n          <p>Start with the words. An agent passes the message on, or acts on it, exactly as written.</p>\n          <table>\n            <thead><tr><th>Unhelpful</th><th>Helpful</th></tr></thead>\n            <tbody>\n              <tr><td>Invalid input</td><td>Enter a 5-digit ZIP code, like 10001</td></tr>\n              <tr><td>Error</td><td>This email address is already registered. Log in or reset your password.</td></tr>\n              <tr><td>Promo code not applied</td><td>The code AUTUMN10 expired on September 30</td></tr>\n              <tr><td>Something went wrong</td><td>We couldn't save your address. Please try again in a minute.</td></tr>\n            </tbody>\n          </table>\n          <ul>\n            <li>Name the field and the problem, and give an example of what's accepted.</li>\n            <li>Say whether it's something the shopper can fix, or something on your side that may work later.</li>\n            <li>Don't rely on color, icons or a border alone. If it isn't in text, assume an agent won't see it.</li>\n            <li>Keep what was entered. Clearing the whole form after one mistake makes an agent start again, and some won't.</li>\n          </ul>\n          <h2>Tie each message to its field</h2>\n          <p>An error floating near a form isn't enough. Agents, like screen readers, need to know which field it belongs to. Two attributes do that:</p>\n          <pre><code>&lt;label for=\"zip\"&gt;ZIP code&lt;/label&gt;\n&lt;input id=\"zip\" name=\"zip\" autocomplete=\"postal-code\"\n       aria-invalid=\"true\" aria-describedby=\"zip-error\"&gt;\n&lt;p id=\"zip-error\"&gt;Enter a 5-digit ZIP code, like 10001&lt;/p&gt;</code></pre>\n          <ul>\n            <li><code>aria-invalid=\"true\"</code> marks the field as having a problem. Remove it once the value is fixed.</li>\n            <li><code>aria-describedby</code> links the field to its message, so the message is read as part of the field's description.</li>\n            <li>Put the message in the page as text next to the field, not in a tooltip that only appears on hover.</li>\n            <li>For required fields, use the <code>required</code> attribute, and the right <code>type</code> and <code>autocomplete</code> values, so agents can get it right first time. Labels matter here too: see <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>.</li>\n          </ul>\n          <h2>Summarize errors and announce changes</h2>\n          <p>On a long form, such as checkout or account sign-up, add a short summary at the top when it's submitted with problems. It lists each error as a link to its field, and gets keyboard focus so it's the first thing read.</p>\n          <pre><code>&lt;div role=\"alert\" tabindex=\"-1\" id=\"error-summary\"&gt;\n  &lt;h2&gt;There are 2 problems with your details&lt;/h2&gt;\n  &lt;ul&gt;\n    &lt;li&gt;&lt;a href=\"#zip\"&gt;Enter a 5-digit ZIP code, like 10001&lt;/a&gt;&lt;/li&gt;\n    &lt;li&gt;&lt;a href=\"#phone\"&gt;Enter a phone number with area code&lt;/a&gt;&lt;/li&gt;\n  &lt;/ul&gt;\n&lt;/div&gt;</code></pre>\n          <p>For messages that appear without a page load, such as \"Promo code applied\" or \"Only 2 left, quantity updated\", use a live region. <code>role=\"alert\"</code> is for urgent errors. <code>role=\"status\"</code> or <code>aria-live=\"polite\"</code> is for confirmations. Keep the message on screen until the next action: toasts that fade after three seconds are easy for an agent to miss.</p>\n          <h2>Make error pages say what happened</h2>\n          <p>Errors aren't only in forms. When an agent follows an old link or a product is discontinued, your error page is what it reads. It should explain, and offer a way forward.</p>\n          <ul>\n            <li><strong>Not found pages</strong> should say the page doesn't exist, and include a search box and links to main categories, so an agent can recover.</li>\n            <li><strong>Discontinued products</strong> are better served by a page that says so, with links to alternatives, or a permanent redirect to the closest replacement.</li>\n            <li><strong>Out-of-stock products</strong> should keep their page, show \"Out of stock\" in text, and set <code>availability</code> to <code>OutOfStock</code> in your product data. Don't turn them into errors.</li>\n            <li><strong>Server errors and maintenance</strong> pages should say it's temporary and, if you can, when to try again.</li>\n          </ul>\n          <h2>Send the status code that matches</h2>\n          <p>Agents and crawlers read your HTTP status code before they read a word of the page. If they disagree, the code usually wins. A \"Page not found\" message sent with a <code>200 OK</code> status, often called a soft 404, tells crawlers the page is fine, so the dead page may stay in AI search answers. A real product page that returns an error tells them it's gone.</p>\n          <table>\n            <thead><tr><th>Situation</th><th>Status code</th></tr></thead>\n            <tbody>\n              <tr><td>Page loaded normally, including out-of-stock products</td><td><code>200</code></td></tr>\n              <tr><td>Page moved for good</td><td><code>301</code> or <code>308</code> to the new address</td></tr>\n              <tr><td>Page doesn't exist</td><td><code>404</code></td></tr>\n              <tr><td>Page removed on purpose and not coming back</td><td><code>410</code></td></tr>\n              <tr><td>Too many requests</td><td><code>429</code>, with a <code>Retry-After</code> header</td></tr>\n              <tr><td>Something broke on your side</td><td><code>500</code></td></tr>\n              <tr><td>Down for maintenance or overloaded</td><td><code>503</code>, with a <code>Retry-After</code> header</td></tr>\n            </tbody>\n          </table>\n          <p>Two common mistakes: redirecting every missing page to the home page, which leaves agents on a page that doesn't answer the question, and serving a block page with <code>200</code>, which looks like content. The meanings of these codes are defined in <a href=\"https://www.rfc-editor.org/rfc/rfc9110\">RFC 9110</a>. For 429s and agents, see <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a>.</p>\n          <p>For form submissions, what matters most is that the page sent back shows the errors in text, tied to their fields. Whether you re-render the form with <code>200</code> or use <code>422</code> is less important than making the message clear.</p>\n          <h2>What AgentScore and Ghost Agents show</h2>\n          <p>AgentScore doesn't have a check for error messages, since it never submits your forms. Several checks cover the groundwork. \"Form fields are labelled\" checks the home page's fields, and \"Cart and checkout fields are labelled for agents\" checks the cart and checkout. \"AI agents can open your product, pricing and cart pages\" compares the status codes AI assistants get with what a browser gets, which catches bot protection sending agents errors that people never see.</p>\n          <p>To see what happens when an agent actually makes a mistake, run a journey. A <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> types the test details you give it and records each step, so if a form rejects them without saying why, the replay shows exactly where it stalled. Ghost Agents stop at the payment form, so payment errors are out of scope. To find the errors agents hit in real traffic, see <a href=\"https://ghostagentlab.com/articles/agent-errors-in-logs/\">finding the errors AI agents hit in your logs</a>.</p>\n          <h2>A quick checklist</h2>\n          <ol>\n            <li>Every error message says what's wrong and how to fix it, in text.</li>\n            <li>Each field with an error has <code>aria-invalid=\"true\"</code> and <code>aria-describedby</code> pointing to its message.</li>\n            <li>Long forms show an error summary that links to each field.</li>\n            <li>Messages that appear without a page load use a live region and stay on screen.</li>\n            <li>Entered values are kept after an error.</li>\n            <li>Missing pages return <code>404</code> or <code>410</code>, with a search box and useful links.</li>\n            <li>Out-of-stock products keep their page, with a <code>200</code> status.</li>\n            <li>Maintenance and rate-limit responses use <code>503</code> or <code>429</code> with <code>Retry-After</code>.</li>\n          </ol>\n          <div>\n            <p><strong>Try it yourself.</strong> Submit your checkout's address form with a wrong ZIP code, then read only the text on the page. If you can't tell which field is wrong and what it wants, neither can an agent.</p>\n          </div>\n          <p>For more on the forms themselves, see <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">how to make checkout work for AI shopping agents</a> and <a href=\"https://ghostagentlab.com/articles/agent-friendly-sign-up-booking/\">sign-up, booking and lead forms AI agents can finish</a>.</p>",
      "date_published": "2026-10-09T13:04:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Navigability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/cart-ai-agents/",
      "url": "https://ghostagentlab.com/articles/cart-ai-agents/",
      "title": "Carts AI agents can manage: quantities, promo codes and totals",
      "summary": "How to build a cart AI agents can change and check: named quantity and remove controls, usable promo code fields, clear totals and carts that last.",
      "content_html": "<p>The cart is where an AI agent checks its work. A person asks for \"two bags of the decaf, and use my code SPRING10\", and the agent has to change a quantity, apply a code and read back a total the person will agree to pay. If the cart makes any of that hard, the agent guesses or gives up. This guide covers what a cart needs so agents can manage it reliably.</p>\n          <h2>Why the cart matters more for agents</h2>\n          <p>A person glances at a cart and knows whether it looks right. An agent has to read it. Before it hands over for payment, a careful agent reports back: what's in the cart, how many, what it costs and when it arrives. It gets every one of those facts from your cart page. If the cart can't be read, the agent can't confirm, and the order stalls at the point where you've already done the hard work of winning it.</p>\n          <p>For the whole journey from search to payment, see <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">how to make checkout work for AI shopping agents</a>. This guide zooms in on the cart itself.</p>\n          <h2>Quantity controls</h2>\n          <p>The most common problem is a pair of \"−\" and \"+\" icons with no names, next to a number that's styled text rather than a field. A person sees a stepper. An agent sees two unnamed buttons and a number it can't change directly.</p>\n          <ul>\n            <li><strong>Use a labelled number field.</strong> A real <code>&lt;input type=\"number\"&gt;</code> whose label names the item lets an agent type the quantity it wants.</li>\n            <li><strong>Name the stepper buttons.</strong> If you keep plus and minus buttons, give each a name: \"Increase quantity of Ethiopia Yirgacheffe\".</li>\n            <li><strong>Say limits in words.</strong> \"Limit 4 per order\", next to the field. An agent that asks for 6 and silently gets 4 will tell the person the wrong thing.</li>\n            <li><strong>Confirm every change.</strong> Update the line total and the cart total on the page, and announce it in a status message, so the agent knows the change took.</li>\n            <li><strong>Make \"Update cart\" a real button</strong> if your cart needs one, placed next to the quantities, not hidden at the bottom.</li>\n          </ul>\n          <p>A cart line that does all of this:</p>\n          <pre><code>&lt;li&gt;\n  &lt;a href=\"/products/ethiopia-yirgacheffe?variant=340g\"&gt;Ethiopia Yirgacheffe, 340 g&lt;/a&gt;\n  &lt;label for=\"qty-1\"&gt;Quantity&lt;/label&gt;\n  &lt;input id=\"qty-1\" name=\"qty-1\" type=\"number\" min=\"1\" max=\"4\" value=\"2\"&gt;\n  &lt;span&gt;Limit 4 per order&lt;/span&gt;\n  &lt;span&gt;$36.00&lt;/span&gt;\n  &lt;button type=\"button\" aria-label=\"Remove Ethiopia Yirgacheffe, 340 g\"&gt;\n    &lt;svg aria-hidden=\"true\"&gt;…&lt;/svg&gt;\n  &lt;/button&gt;\n&lt;/li&gt;\n&lt;p role=\"status\"&gt;Quantity updated to 2. Subtotal $36.00.&lt;/p&gt;</code></pre>\n          <h2>Remove buttons</h2>\n          <ul>\n            <li><strong>Name what each one removes.</strong> On a cart with three items, three buttons called \"×\", or nothing at all, leave the agent guessing which is which. \"Remove Ethiopia Yirgacheffe, 340 g\" leaves no doubt.</li>\n            <li><strong>Confirm in text.</strong> \"Ethiopia Yirgacheffe removed. Undo.\" A line that simply vanishes could mean it was removed, or that the page is still loading.</li>\n            <li><strong>Don't make zero the only way out.</strong> Setting the quantity to 0 should work, but a clear remove button shouldn't be missing.</li>\n          </ul>\n          <p>AgentScore's \"Buttons and links have names\" check looks for exactly this kind of icon-only control on your home page. Apply the same rule in the cart, where it matters more.</p>\n          <h2>Promo code fields</h2>\n          <p>People often ask an assistant to use a code they already have. That only works if the agent can find the field, apply the code and tell whether it worked.</p>\n          <ul>\n            <li><strong>Give the field a visible label</strong>, such as \"Discount code\", not just placeholder text. Placeholders vanish as soon as typing starts, and many agents ignore them.</li>\n            <li><strong>Pair it with a real \"Apply\" button.</strong> Not a form that only submits when someone presses Enter.</li>\n            <li><strong>If it's collapsed behind \"Have a code?\"</strong>, make that a real button that says it opens something (<code>aria-expanded</code>), not a styled line of text.</li>\n            <li><strong>Explain failures in words.</strong> \"SPRING10 expired on March 31\" or \"This code doesn't apply to sale items\" lets the agent tell the person why. \"Invalid code\" in red does not. See <a href=\"https://ghostagentlab.com/articles/error-messages-ai-agents/\">error messages AI agents can understand</a>.</li>\n            <li><strong>Show the applied discount as its own line</strong>, with the code name, the amount and a named button to remove it.</li>\n          </ul>\n          <h2>Totals the agent can repeat back</h2>\n          <p>The goal is simple: an agent should be able to say \"The total is $44.18, including $5.00 shipping and $3.18 tax\" and be right. That means each part of the total is on the page as text.</p>\n          <table>\n            <thead><tr><th>Line</th><th>What to show</th></tr></thead>\n            <tbody>\n              <tr><td>Subtotal</td><td>The sum of the items, as text, updated after every change</td></tr>\n              <tr><td>Discounts</td><td>Each code or promotion by name, with its amount</td></tr>\n              <tr><td>Shipping</td><td>The cost for the chosen method, or an estimate with the method named. If it truly depends on the address, say \"Calculated at checkout\"</td></tr>\n              <tr><td>Tax</td><td>The amount, or \"Calculated at checkout\". Say whether prices already include tax</td></tr>\n              <tr><td>Total</td><td>The final figure with its currency</td></tr>\n            </tbody>\n          </table>\n          <ul>\n            <li>Show the currency clearly. \"$\" alone is ambiguous on a site that sells in several countries. See <a href=\"https://ghostagentlab.com/articles/international-sites-ai-agents/\">international sites for AI agents</a>.</li>\n            <li>Put free-shipping thresholds in words: \"Spend $12.00 more for free shipping.\" A progress bar alone tells an agent nothing.</li>\n            <li>Keep prices as text in the page, not in images or drawn on a canvas. See <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">prices AI agents can read</a>.</li>\n          </ul>\n          <h2>Carts that survive the handoff</h2>\n          <p>Most agents hand back to the person to sign in or pay, and that can happen hours later, or on another device. A cart that empties in the meantime undoes the agent's work.</p>\n          <ul>\n            <li><strong>Keep carts for days, not minutes.</strong> Don't clear them on a short session timeout.</li>\n            <li><strong>Keep the cart at a stable URL</strong>, such as <code>/cart</code>, as a full page as well as any slide-out drawer.</li>\n            <li><strong>Consider cart links.</strong> Some platforms can rebuild a cart from a URL (Shopify calls these cart permalinks), which lets an agent hand the person a link that opens the same cart. Ask your platform what it supports.</li>\n            <li><strong>Explain changes.</strong> If stock or a price changes while the cart waits, say so on the page: \"Only 1 left. We've changed your quantity to 1.\"</li>\n          </ul>\n          <h2>Upsell pop-ups and cart drawers</h2>\n          <p>An upsell window after \"Add to cart\", or on the way to checkout, is one more thing an agent has to get past.</p>\n          <ul>\n            <li>If you use one, make it a proper dialog with a heading, a \"No thanks\" button with text, and a named close button. See <a href=\"https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/\">cookie banners and pop-ups</a>.</li>\n            <li><strong>Never pre-add extras</strong>, such as a pre-ticked protection plan. An agent may not notice, and the person pays for something they didn't ask for. In some markets, such as the EU, consumer law already restricts pre-ticked extras.</li>\n            <li>Recommendations under the cart are fine. Keep the checkout button above them.</li>\n            <li>Give a cart drawer a labelled close button and a real link to the full cart page.</li>\n          </ul>\n          <h2>What AgentScore checks on your cart</h2>\n          <p>The free <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> scan is read-only: it loads pages but never adds anything to a cart. On many stores the cart is empty when it looks, and the scan says so rather than guessing. For stores, these checks cover the cart:</p>\n          <ul>\n            <li><strong>AI agents can open your product, pricing and cart pages.</strong> It requests your cart as ChatGPT-User, Claude-User and Perplexity-User and compares what they get with what a browser gets.</li>\n            <li><strong>AI agents aren't blocked or shown a CAPTCHA at the cart and checkout.</strong></li>\n            <li><strong>Cart and checkout fields are labelled for agents.</strong> Labels rather than placeholders, the right input types, and autocomplete. If the empty cart shows no fields, this is marked \"not tested\".</li>\n            <li><strong>Shoppers can check out without an account.</strong> It reads the cart for wording like \"Sign in to check out\" or \"Check out as a guest\".</li>\n            <li><strong>Key pages are linked from the home page</strong>, including the cart, with ordinary links in the HTML.</li>\n          </ul>\n          <p>Changing quantities, applying codes and reading totals need items in the cart, so they belong to Task completion, which the free scan marks \"not tested\". To test them, set up a <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> test with a goal in your own words, such as \"Add two bags of decaf, apply the code TEST10 and go to checkout\", using a test code you've created, and a pass condition that looks for the code's name on the final page. Ghost Agents stop at the payment form, so no order is placed. See <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">how to test your key journeys with AI agents</a>.</p>\n          <div>\n            <p><strong>A five-minute cart test.</strong> Add two items, then use only the keyboard to change a quantity, remove an item and apply a code. Then read the cart aloud as if over the phone: items, quantities, discount, shipping, tax, total. Anything you had to look at a picture to work out, an agent can't see.</p>\n          </div>",
      "date_published": "2026-10-09T13:03:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Task completion"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/order-tasks-ai-agents/",
      "url": "https://ghostagentlab.com/articles/order-tasks-ai-agents/",
      "title": "Order tracking, returns and reorders by AI agents",
      "summary": "How to let AI agents track orders, start returns and reorder for customers without a login, and when to hand back to the person for security.",
      "content_html": "<p>Buying is only part of what people ask AI assistants to do. \"Where's my order?\" \"Return the shoes, they're too small.\" \"Order the same coffee again.\" These after-sale tasks are frequent, they cost you nothing when customers can do them on their own, and they're often locked behind a login an agent can't use. This guide covers how to make them work for AI agents, and where the agent should hand back to the person.</p>\n          <h2>Why after-sale tasks matter</h2>\n          <p>\"Where is my order?\" is one of the most common reasons customers contact an online store. When an assistant can answer it from your site, that's a support ticket that never gets written. When it can't, the person contacts you anyway, or the assistant pieces together an answer from a carrier's site and gets it wrong.</p>\n          <p>Returns matter before the sale, too. Assistants read return policies when comparing stores, and a clear, easy returns process is part of what makes a store safe to recommend. See <a href=\"https://ghostagentlab.com/articles/policy-pages-ai-assistants/\">policy pages AI assistants can quote</a>.</p>\n          <h2>Order tracking without a login</h2>\n          <p>An agent can't sign in as the person, but it can usually be given an order number and an email address. That's enough for a well-built order lookup.</p>\n          <ul>\n            <li><strong>Give order lookup its own page</strong> at a stable URL, such as <code>/orders/lookup</code>, linked from your footer and help pages with an ordinary link.</li>\n            <li><strong>Ask for two things:</strong> the order number and the email address used at checkout. Label both fields, and say where to find the number: \"It starts with NW- and is in your confirmation email.\"</li>\n            <li><strong>Show the status in words.</strong> \"Shipped on October 3 with UPS. Expected Tuesday, October 7\", the tracking number as a link, and the items in the order. A progress bar alone tells an agent nothing.</li>\n            <li><strong>Put a direct link in every order email</strong> that opens the order's status page without signing in. Many platforms offer one. Agents reading an email for the person can follow it straight to the answer.</li>\n            <li><strong>Keep failures helpful but safe.</strong> \"We couldn't find an order matching those details\" is right. Don't say which of the two was wrong, which would let anyone test email addresses.</li>\n            <li><strong>Limit guessing with rate limits, not a CAPTCHA on every lookup.</strong> Answer too many attempts with a 429 status. See <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a>.</li>\n          </ul>\n          <p>A lookup form agents can fill in:</p>\n          <pre><code>&lt;form action=\"/orders/lookup\" method=\"post\"&gt;\n  &lt;label for=\"order\"&gt;Order number&lt;/label&gt;\n  &lt;input id=\"order\" name=\"order\" autocomplete=\"off\" aria-describedby=\"order-hint\"&gt;\n  &lt;p id=\"order-hint\"&gt;Starts with NW-. You'll find it in your confirmation email.&lt;/p&gt;\n  &lt;label for=\"email\"&gt;Email used at checkout&lt;/label&gt;\n  &lt;input id=\"email\" name=\"email\" type=\"email\" autocomplete=\"email\"&gt;\n  &lt;button type=\"submit\"&gt;Find my order&lt;/button&gt;\n&lt;/form&gt;</code></pre>\n          <p>Show only what's needed on a page reached this way: status, items, carrier and the delivery city. Leave out the full address, phone number and payment details.</p>\n          <h2>Returns and exchanges</h2>\n          <ul>\n            <li><strong>Start from a plain-text policy</strong>: how long customers have, what condition items must be in, who pays return shipping, how long refunds take and what can't be returned.</li>\n            <li><strong>Open the returns portal with an order number and email</strong>, the same as tracking. If returns need an account, most agents stop at the door.</li>\n            <li><strong>Label every choice.</strong> Items as checkboxes named with the product and size, the reason as a list of options in words, and refund, exchange or store credit as radio buttons that say what each one means: \"Store credit: issued as soon as we receive the item.\"</li>\n            <li><strong>Explain why an item isn't eligible.</strong> \"This item was delivered 45 days ago. Returns close after 30 days.\" See <a href=\"https://ghostagentlab.com/articles/error-messages-ai-agents/\">error messages AI agents can understand</a>.</li>\n            <li><strong>Summarize before submitting.</strong> A review step (\"You're returning 1 item for a $64.00 refund to your original payment method\") lets the agent check with the person first.</li>\n            <li><strong>Say what happens next in words.</strong> Where the label is (a link to the PDF, and \"We've emailed it to you\"), drop-off options, and when the refund will arrive.</li>\n          </ul>\n          <h2>Reorders</h2>\n          <p>\"Order the same again\" usually needs order history, and order history usually needs an account. You can still make it easy:</p>\n          <ul>\n            <li><strong>Add a \"Buy again\" button to the order status page</strong>, the one reached from the email link. Name it after the item: \"Buy again: Ethiopia Yirgacheffe, 340 g\".</li>\n            <li><strong>Put the same item back in the cart</strong>: the same size, color and quantity. Say if anything changed: \"The price is now $19.00\" or \"This size is out of stock.\"</li>\n            <li><strong>Keep product URLs stable.</strong> A product link in a year-old confirmation email should still work, or redirect to the replacement product with a note saying so. See <a href=\"https://ghostagentlab.com/articles/canonical-urls-ai-agents/\">canonical URLs and duplicate pages</a>.</li>\n            <li><strong>Make subscription controls clear</strong> if you sell subscriptions. \"Skip next delivery\" and \"Change quantity\" as named buttons, with the next delivery date in words.</li>\n          </ul>\n          <h2>Security, accounts and handing back to the person</h2>\n          <p>These tasks touch personal details and money, so there's a real tension. Agents shouldn't need the person's password, and anyone holding an order number shouldn't be able to change an order. The answer is to match the check to the risk:</p>\n          <table>\n            <thead><tr><th>Task</th><th>What it should take</th><th>Who finishes it</th></tr></thead>\n            <tbody>\n              <tr><td>Track an order</td><td>Order number and email</td><td>The agent</td></tr>\n              <tr><td>Check return eligibility and options</td><td>Order number and email</td><td>The agent</td></tr>\n              <tr><td>Submit a return or exchange</td><td>Order number and email, plus a review step</td><td>The agent, once the person agrees</td></tr>\n              <tr><td>Reorder</td><td>A \"Buy again\" link or button</td><td>The agent builds the cart, the person pays</td></tr>\n              <tr><td>Change a delivery address or cancel</td><td>A one-time code or emailed link</td><td>The person</td></tr>\n              <tr><td>Change email, password or saved cards</td><td>Full sign-in, ideally with two-step sign-in</td><td>The person only</td></tr>\n            </tbody>\n          </table>\n          <ul>\n            <li><strong>Use step-up checks for risky actions only.</strong> A one-time code sent by email or text is a clean handoff point: the agent stops and says \"Northwind has sent a code to your email\", and the person takes it from there.</li>\n            <li><strong>Make the handoff obvious.</strong> Browser agents typically hand control back for sign-in and payment. A sign-in form with labelled fields, and support for passkeys or password managers, makes that moment quick for the person.</li>\n            <li><strong>Never ask for card details to prove who someone is.</strong> No one, human or agent, should be typing card numbers to see an order status.</li>\n            <li><strong>Keep help content public.</strong> Delivery times, return policy and FAQs don't need a login. See <a href=\"https://ghostagentlab.com/articles/login-walls-ai-agents/\">login walls and gated content</a>.</li>\n          </ul>\n          <p>The same thinking applies to buying: see <a href=\"https://ghostagentlab.com/articles/guest-checkout-ai-agents/\">guest checkout</a> and <a href=\"https://ghostagentlab.com/articles/agent-friendly-sign-up-booking/\">sign-up and booking forms agents can use</a>.</p>\n          <h2>How to test these journeys</h2>\n          <ol>\n            <li><strong>Try it yourself with an assistant.</strong> Place a test order, then ask an assistant that can browse to find its status using the order number and email. Watch where it stalls.</li>\n            <li><strong>Check the basics with <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>.</strong> It checks that AI assistants are let in and can open your key product, pricing and cart pages. It doesn't test order lookup or returns, and the free scan marks Task completion \"not tested\".</li>\n            <li><strong>Monitor them with <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a>.</strong> Describe the journey in your own words, such as \"Find the order tracking page, look up order NW-10422 with test@northwind.example, and report the delivery status\", and set a pass condition that looks for text like \"Shipped\" on the final page. For returns, write the goal so the agent stops at the review step, without submitting.</li>\n          </ol>\n          <div>\n            <p><strong>Start with tracking.</strong> It's the most common after-sale question, it needs no new security thinking, and a public lookup page with a direct link in every order email works for customers and their agents alike.</p>\n          </div>",
      "date_published": "2026-10-09T13:02:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Task completion"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/agent-readiness-kpis/",
      "url": "https://ghostagentlab.com/articles/agent-readiness-kpis/",
      "title": "Agent readiness KPIs: what to report each month",
      "summary": "A monthly agent readiness report in six measures: AgentScore, agent visits, verification, blocks and errors, Ghost Agent pass rates and AI referrals.",
      "content_html": "<p>Agent readiness work needs a monthly number, or it quietly slides down the priority list. This guide sets out a one-page monthly report built on six measures: where each one comes from, what good looks like, and how to read them together. There are no reliable industry benchmarks for most of these yet, so the comparison that matters is your own trend.</p>\n          <h2>Four rules for the report</h2>\n          <ul>\n            <li><strong>Trend over level.</strong> A score of 68 means little on its own. A score that went from 52 to 68 in a quarter, with the reasons, means a lot.</li>\n            <li><strong>Keep two questions apart.</strong> \"Can AI agents use our site?\" (readiness, errors, test results) and \"Is it turning into business?\" (referrals and orders). The first leads; the second follows.</li>\n            <li><strong>One owner per measure.</strong> Each number needs someone who explains it and acts on it. See <a href=\"https://ghostagentlab.com/blog/who-owns-agent-readiness/\">who owns agent readiness</a>.</li>\n            <li><strong>Note what changed.</strong> Releases, bot protection changes, new apps and campaigns. Most movements in these numbers trace back to one of them.</li>\n          </ul>\n          <h2>The six measures</h2>\n          <table>\n            <thead><tr><th>Measure</th><th>Question it answers</th><th>Where it comes from</th></tr></thead>\n            <tbody>\n              <tr><td>AgentScore</td><td>Can agents get in, read and find their way around?</td><td>An AgentScore scan</td></tr>\n              <tr><td>Agent visits by type</td><td>Who's visiting, and what are they reading?</td><td>Server or CDN logs</td></tr>\n              <tr><td>Verification share</td><td>How much of that traffic is really who it says it is?</td><td>Logs checked against operators' published proof</td></tr>\n              <tr><td>Blocks and errors</td><td>Are we turning agents away?</td><td>Status codes in your logs</td></tr>\n              <tr><td>Ghost Agent pass rates</td><td>Can an agent actually finish the job?</td><td>Ghost Agent tests</td></tr>\n              <tr><td>AI referral visits and orders</td><td>Is it bringing in customers and revenue?</td><td>Your analytics and your store's orders</td></tr>\n            </tbody>\n          </table>\n          <p>In Ghost Agent Labs, the site's Overview page brings several of these together. The sections below name the page that holds each one.</p>\n          <h2>1. AgentScore</h2>\n          <p><strong>Report:</strong> the overall score and its band, the scores for Access, Readability and Navigability, and how many findings you fixed and how many are new. The free scan doesn't test Task completion yet, so the score covers the other three categories.</p>\n          <p><strong>Where:</strong> the Readiness page. Every scan is saved, so you can track the score over time, and \"All findings\" exports to CSV for your developers.</p>\n          <p><strong>Good looks like:</strong> a steady or rising score, no critical findings open for more than a month, and every drop explained. Scan on the same day each month and after major releases, so the numbers compare like with like. The bands are 80 and up Good, 60 to 79 Fair, 40 to 59 Weak and below 40 Poor. See <a href=\"https://ghostagentlab.com/blog/reading-your-agentscore-report/\">reading your AgentScore report</a>.</p>\n          <h2>2. Agent visits by type</h2>\n          <p><strong>Report:</strong> requests by visitor type (search crawlers, AI training crawlers, AI assistants and autonomous agents) compared with last month, the pages agents read most, and any new agents.</p>\n          <p><strong>Where:</strong> the Agent traffic page (\"Requests by visitor type\", \"Agent sessions\" and \"Pages agents request most\"), and \"What changed\" on the Overview, which lists first-time agents and agents whose traffic rose or fell by half or more.</p>\n          <p><strong>Good looks like:</strong> AI assistant and AI search traffic holding steady or growing, and agents reaching product, pricing and policy pages rather than stalling on the home page or old URLs. Don't celebrate total volume. A busy training crawler isn't a customer. Assistant requests are the closest thing to a customer visit from AI, because each one is usually a person's question. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure AI agent traffic</a>.</p>\n          <h2>3. Verification share</h2>\n          <p><strong>Report:</strong> of the requests claiming to be a known agent, the share that were verified, couldn't be verified, and were spoofed.</p>\n          <p><strong>Where:</strong> the Verification &amp; spoofing page.</p>\n          <p><strong>Good looks like:</strong> spoofed traffic close to zero, or falling after you block it, and a verified share that rises as more operators publish ways to check their agents. This measure also keeps the rest of the report honest: never report \"ChatGPT visited 40,000 times\" from user agent names alone, because impostors use those names too. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</p>\n          <h2>4. Blocks and errors</h2>\n          <p><strong>Report:</strong> the share of agent requests that were blocked (401, 403 and 429 responses), the 404 and server error rates, and the agents and pages most affected.</p>\n          <p><strong>Where:</strong> \"Errors: agents vs people\" on the Overview, \"What agents got back\" on the Agent traffic page, and the status codes on each agent's page.</p>\n          <p><strong>Good looks like:</strong> blocked responses to verified AI assistants near zero, agents getting errors about as often as people do rather than much more, and agent 404s trending down as you redirect old URLs. A sudden rise almost always follows a change to bot protection or rate limits. See <a href=\"https://ghostagentlab.com/articles/agent-errors-in-logs/\">finding the errors AI agents hit</a>.</p>\n          <h2>5. Ghost Agent pass rates</h2>\n          <p><strong>Report:</strong> for each journey you test, such as checkout, product search or sign-up, whether it's passing, flaky or failing, its pass rate, and the step where failures cluster.</p>\n          <p><strong>Where:</strong> the Ghost Agents page, with a replay of every run.</p>\n          <p><strong>Good looks like:</strong> your money journeys passing, flaky results treated as findings rather than noise, and failures fixed within a sprint. A pass rate below 100% means some customers' assistants would fail too. This is the measure closest to revenue, because it tests whether an agent can do the job, not just whether the page looks right. See <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">how to test your key journeys with AI agents</a>.</p>\n          <h2>6. AI referral visits and orders</h2>\n          <p><strong>Report:</strong> visitors who arrived from AI assistants' answers, by assistant, their conversion rate, and orders and revenue from AI.</p>\n          <p><strong>Where:</strong> your analytics, set up as in <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">tracking visits and sales from AI assistants</a>. In Ghost Agent Labs, \"Visitors from AI assistants\" is on the Overview. Send your store's orders to the site's orders URL and Mission Control shows \"Agent-driven revenue\": orders referred by an AI answer or placed by an agent, and their share of revenue.</p>\n          <p><strong>Good looks like:</strong> growth from whatever base you start at, and conversion from AI referrals that holds up against your other referral sources. Present it as a floor, not a total: app traffic and copied links arrive as \"direct\", and many AI answers influence a purchase without a click.</p>\n          <h2>Reading the measures together</h2>\n          <p>Single numbers mislead. These combinations are where the report earns its keep:</p>\n          <ul>\n            <li><strong>Score steady, blocks up.</strong> Something changed in bot protection or rate limits. Check the agent's page for which pages are affected and when it started.</li>\n            <li><strong>Visits up, referrals flat.</strong> Assistants are reading you but not sending people. Look at product data, prices and policy pages: are they giving assistants what they need to recommend you?</li>\n            <li><strong>Score unchanged, pass rate down.</strong> A journey broke in a way the scan can't see: a new pop-up, a redesigned size picker, a checkout app. The run replay shows where.</li>\n            <li><strong>Referrals up, checkout test failing.</strong> AI is sending you customers and losing some of them at the last step. This is the most urgent combination on the page.</li>\n            <li><strong>Spoofed share up.</strong> Scrapers are using agent names. It doesn't hurt readiness, but it inflates traffic numbers and costs bandwidth.</li>\n          </ul>\n          <h2>A one-page template</h2>\n          <ol>\n            <li><strong>Headline.</strong> One sentence: better, worse or the same, and why.</li>\n            <li><strong>Scorecard.</strong> The six measures, this month against last month, with an arrow and a short note each.</li>\n            <li><strong>What changed.</strong> Site releases, bot protection changes and new agents, linked to the numbers they moved.</li>\n            <li><strong>Fixed this month.</strong> The findings closed and the journeys repaired.</li>\n            <li><strong>Next month.</strong> The top three fixes, each with an owner.</li>\n            <li><strong>Decisions needed.</strong> For example, whether to allow a new AI agent, or budget for platform work.</li>\n          </ol>\n          <div>\n            <p><strong>What to leave out.</strong> Total bot requests without a breakdown, scores without a trend, and industry comparisons nobody can source. They invite the wrong questions. If you're presenting upward, see <a href=\"https://ghostagentlab.com/blog/explain-agent-readiness-to-your-board/\">how to explain agent readiness to your board</a>, and for what to fix first, <a href=\"https://ghostagentlab.com/blog/90-day-agent-readiness-plan/\">the 90-day agent readiness plan</a>.</p>\n          </div>",
      "date_published": "2026-10-09T13:01:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Measurement"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/agent-errors-in-logs/",
      "url": "https://ghostagentlab.com/articles/agent-errors-in-logs/",
      "title": "Finding the errors AI agents hit, in your server and CDN logs",
      "summary": "How to find the 403s, 429s, challenges, timeouts and server errors AI agents hit in your logs, with example queries and how to separate real agents from fakes.",
      "content_html": "<p>When an AI agent can't get into your site, nobody tells you. There's no complaint and no support ticket. The assistant just answers with someone else's products. The evidence is in your server and CDN logs. This guide shows which errors to look for, how to find them with a few queries, and how to separate real agents from impostors before you change anything.</p>\n          <p>It builds on <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure AI agent traffic</a>, which covers getting hold of your logs in the first place.</p>\n          <h2>The errors that matter</h2>\n          <table>\n            <thead><tr><th>What you see</th><th>What it usually means for an agent</th><th>Who usually fixes it</th></tr></thead>\n            <tbody>\n              <tr><td>403 Forbidden</td><td>Bot protection, a firewall rule or a country block turned it away</td><td>Bot protection or CDN admin</td></tr>\n              <tr><td>401 Unauthorized</td><td>The page needs a login</td><td>E-commerce or product team</td></tr>\n              <tr><td>429 Too Many Requests</td><td>It hit a rate limit</td><td>CDN admin or developers</td></tr>\n              <tr><td>Challenge page</td><td>A \"checking your browser\" page or CAPTCHA it can't solve. Often a 403 or 503, sometimes a 200</td><td>Bot protection or CDN admin</td></tr>\n              <tr><td>500, 502, 503</td><td>Your server failed. Agents can reach pages people rarely visit, which may not be cached</td><td>Developers or hosting</td></tr>\n              <tr><td>504, slow responses, 499</td><td>The page took too long. In nginx, 499 means the client gave up and closed the connection first</td><td>Developers or hosting</td></tr>\n              <tr><td>404 Not Found</td><td>An old or guessed URL. AI models often remember URLs you've since changed</td><td>Content or SEO team</td></tr>\n            </tbody>\n          </table>\n          <p>Challenge pages are the easiest to miss, because some bot protection serves them with a 200 status, so they look like successes. Look for your CDN's security or bot action field instead. Cloudflare, for example, adds a <code>cf-mitigated: challenge</code> header to challenge responses. A response that's much smaller than the real page usually is one too.</p>\n          <h2>What your logs need</h2>\n          <p>Most logs already have the basics. Check yours include:</p>\n          <ul>\n            <li><strong>Time, client IP address and user agent.</strong> The IP address is what lets you verify the agent later.</li>\n            <li><strong>Path and status code.</strong> Query strings can be dropped; you rarely need them, and they can hold personal data.</li>\n            <li><strong>Response size and response time.</strong> These catch challenge pages and timeouts.</li>\n            <li><strong>The CDN's security action</strong>, if it has one: allowed, blocked, challenged, rate limited.</li>\n          </ul>\n          <p>On nginx, a log format like this captures them:</p>\n          <pre><code>log_format agents '$time_iso8601 $remote_addr \"$request_method $uri\" '\n                  '$status $body_bytes_sent $request_time \"$http_user_agent\"';</code></pre>\n          <p>If you only have the standard \"combined\" access log, this one-liner counts status codes for three AI assistants:</p>\n          <pre><code>grep -E 'ChatGPT-User|Claude-User|Perplexity-User' access.log \\\n  | awk '{print $9}' | sort | uniq -c | sort -rn</code></pre>\n          <h2>Example queries</h2>\n          <p>Once logs are in a database or log tool, a few queries answer most questions. The examples below are illustrative, written for PostgreSQL against a table called <code>requests</code> with columns <code>ts</code>, <code>client_ip</code>, <code>user_agent</code>, <code>path</code>, <code>status</code>, <code>bytes</code>, <code>duration_ms</code> and <code>security_action</code>. Rename them to match your own log store; the same ideas work in BigQuery, Athena or a log search tool.</p>\n          <h3>Error rates by agent, last 7 days</h3>\n          <pre><code>SELECT\n  CASE\n    WHEN user_agent ILIKE '%ChatGPT-User%'    THEN 'ChatGPT-User'\n    WHEN user_agent ILIKE '%OAI-SearchBot%'   THEN 'OAI-SearchBot'\n    WHEN user_agent ILIKE '%Claude-User%'     THEN 'Claude-User'\n    WHEN user_agent ILIKE '%Perplexity-User%' THEN 'Perplexity-User'\n    WHEN user_agent ILIKE '%PerplexityBot%'   THEN 'PerplexityBot'\n    WHEN user_agent ILIKE '%Googlebot%'       THEN 'Googlebot'\n    ELSE 'everything else'\n  END AS agent,\n  COUNT(*) AS requests,\n  COUNT(*) FILTER (WHERE status IN (401, 403)) AS blocked,\n  COUNT(*) FILTER (WHERE status = 429)         AS rate_limited,\n  COUNT(*) FILTER (WHERE status = 404)         AS not_found,\n  COUNT(*) FILTER (WHERE status &gt;= 500)        AS server_errors,\n  ROUND(100.0 * COUNT(*) FILTER (WHERE status &gt;= 400) / COUNT(*), 1) AS error_pct\nFROM requests\nWHERE ts &gt;= now() - interval '7 days'\nGROUP BY 1\nORDER BY requests DESC;</code></pre>\n          <p>\"Everything else\" mixes people with other bots, many of them scanners probing for pages that don't exist. For a fair baseline, compare with requests from ordinary browsers: if AI agents get errors far more often than people do, something is turning them away.</p>\n          <h3>Where AI assistants are turned away</h3>\n          <pre><code>SELECT path, status, COUNT(*) AS hits\nFROM requests\nWHERE ts &gt;= now() - interval '7 days'\n  AND user_agent ~* '(ChatGPT-User|Claude-User|Perplexity-User)'\n  AND (status IN (401, 403, 429) OR security_action = 'challenge')\nGROUP BY path, status\nORDER BY hits DESC\nLIMIT 20;</code></pre>\n          <h3>Challenges and timeouts by day</h3>\n          <pre><code>SELECT date_trunc('day', ts) AS day,\n  COUNT(*) FILTER (WHERE security_action = 'challenge')   AS challenged,\n  COUNT(*) FILTER (WHERE status IN (499, 504))            AS timed_out,\n  COUNT(*) FILTER (WHERE duration_ms &gt; 10000)             AS over_10s\nFROM requests\nWHERE user_agent ~* '(ChatGPT-User|Claude-User|Perplexity-User|OAI-SearchBot)'\nGROUP BY 1\nORDER BY 1;</code></pre>\n          <p>A step change on one day usually lines up with a release or a bot protection change. That's your first lead.</p>\n          <h2>Separating verified agents from claimed ones</h2>\n          <p>Everything above trusts the user agent, and a user agent is only a claim. Scrapers often call themselves ChatGPT-User or Googlebot to get past bot protection. When your firewall blocks them, that 403 is working as intended. So before you loosen any rule, split agent errors into three groups:</p>\n          <ul>\n            <li><strong>Verified.</strong> The request came from the operator's published IP ranges, passed forward-confirmed reverse DNS, or carried a valid signature (Web Bot Auth, still an emerging standard). Errors here are real problems to fix.</li>\n            <li><strong>Spoofed.</strong> It used a known agent's name but failed the operator's check. Blocking these is correct.</li>\n            <li><strong>Unverifiable.</strong> The operator doesn't publish a way to check. Decide case by case.</li>\n          </ul>\n          <p>Operators such as OpenAI publish their agents' IP ranges (see <a href=\"https://platform.openai.com/docs/bots\">OpenAI's crawler documentation</a>). Load them into a table, refresh it regularly, and join:</p>\n          <pre><code>-- agent_ranges(agent text, cidr cidr), loaded from each operator's published list\nSELECT r.agent IS NOT NULL AS verified, l.status, COUNT(*) AS hits\nFROM requests l\nLEFT JOIN agent_ranges r\n  ON l.client_ip::inet &lt;&lt; r.cidr\n AND l.user_agent ILIKE '%' || r.agent || '%'\nWHERE l.user_agent ~* '(ChatGPT-User|OAI-SearchBot)'\n  AND l.ts &gt;= now() - interval '7 days'\nGROUP BY 1, 2\nORDER BY 1 DESC, hits DESC;</code></pre>\n          <p>For the methods in detail, including reverse DNS and signed requests, see <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</p>\n          <div>\n            <p><strong>Mind the privacy.</strong> Logs hold IP addresses, which count as personal data for human visitors. Filter to agent user agents before you keep detailed records, hash or drop IP addresses for everyone else, and keep a retention period that matches your privacy policy.</p>\n          </div>\n          <h2>From errors to fixes</h2>\n          <ul>\n            <li><strong>Verified assistants get 403s on product or policy pages:</strong> allow verified bots in your bot protection. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a> and <a href=\"https://ghostagentlab.com/articles/cdn-settings-ai-agents/\">setting up your CDN for AI agents</a>.</li>\n            <li><strong>Repeated 429s:</strong> give verified agents their own limits, and send <code>Retry-After</code>. See <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a>.</li>\n            <li><strong>Blocks only from certain countries or networks:</strong> see <a href=\"https://ghostagentlab.com/articles/geo-blocking-ai-agents/\">geo-blocking, VPN blocks and AI agents</a>.</li>\n            <li><strong>Challenges at the cart or checkout:</strong> these cost orders directly. AgentScore checks for them, and <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">the checkout guide</a> covers the fix.</li>\n            <li><strong>The same 404s week after week:</strong> redirect old URLs to their current pages.</li>\n            <li><strong>Server errors and timeouts:</strong> cache the pages agents request most, and check them in your monitoring like any other.</li>\n          </ul>\n          <p>If writing queries isn't your team's idea of a good week, Ghost Agent Labs does this from your Cloudflare, Vercel or JSON logs. The Agent traffic page shows what agents got back, each agent's page shows its status codes, the Verification &amp; spoofing page splits verified from spoofed, and Alerts tell you when a major AI assistant starts being blocked. Either way, check after every bot protection or CDN change, and at least monthly as part of your <a href=\"https://ghostagentlab.com/articles/agent-readiness-kpis/\">agent readiness KPIs</a>.</p>",
      "date_published": "2026-10-09T13:00:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Measurement"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/seo-to-agent-readiness/",
      "url": "https://ghostagentlab.com/blog/seo-to-agent-readiness/",
      "title": "SEO got you found. Agent readiness gets you chosen",
      "summary": "SEO decides whether AI agents find you. Agent readiness decides whether they can use your site and choose you. What carries over, and what's new.",
      "content_html": "<p>For twenty years, SEO has answered one question: when someone searches, do they find you? AI agents add a second question that SEO was never designed to answer. When an agent arrives on someone's behalf, can it actually use your site? Getting found is still the start. Getting chosen now depends on what happens next.</p>\n          <h2>What SEO was built for</h2>\n          <p>Search engines crawl, index and rank. A person sees a list of results, clicks one, and does the rest themselves. Everything after the click is the person's job: reading the page, closing the cookie banner, finding the size selector, getting through checkout. If the site is awkward, people usually push through.</p>\n          <p>SEO teams have become very good at the part before the click. Titles, descriptions, sitemaps, structured data, internal links, page speed, crawl budgets. All of that still matters.</p>\n          <h2>What changes when an agent does the clicking</h2>\n          <p>When someone asks an AI assistant to compare three running shoes, check which one is in stock in a size 10, and buy it, there is no results page and often no click by a person. The assistant fetches pages itself. A browser agent may open your site, search your catalog and try to add to cart.</p>\n          <p>That shifts the work from the person to the agent, and agents don't push through. If the price isn't readable, the agent can't quote it. If a pop-up has no labelled close button, the journey ends there. If bot protection challenges the assistant's request, the person gets a competitor's answer instead of yours. Nobody tells you. The visit simply doesn't turn into anything.</p>\n          <p>So the question moves from \"do we rank?\" to \"when an agent arrives, can it get in, understand the page, find its way and finish the job?\" That's what we mean by <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness</a>.</p>\n          <h2>Where SEO already gets you most of the way</h2>\n          <p>The good news for SEO teams is that a lot of the groundwork is shared. Several AgentScore checks will look familiar.</p>\n          <table>\n            <thead><tr><th>SEO practice</th><th>What it does for AI agents</th></tr></thead>\n            <tbody>\n              <tr><td>A clean robots.txt</td><td>The same file controls AI assistants, AI search and training crawlers, each by its own name. See <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>.</td></tr>\n              <tr><td>Clear titles, descriptions and headings</td><td>How an agent decides whether a page answers the question it was sent to ask. See <a href=\"https://ghostagentlab.com/articles/titles-descriptions-headings-ai/\">titles, descriptions and headings AI agents understand</a>.</td></tr>\n              <tr><td>Product structured data</td><td>Lets an assistant quote your price and stock correctly instead of guessing. See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data AI shopping agents can read</a>.</td></tr>\n              <tr><td>A valid XML sitemap</td><td>Helps agents find pages that aren't linked prominently. See <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">XML sitemaps for AI agents</a>.</td></tr>\n              <tr><td>Image alt text</td><td>The only way an agent reading text knows what a product photo shows. See <a href=\"https://ghostagentlab.com/articles/image-alt-text-ai-agents/\">alt text for AI agents</a>.</td></tr>\n            </tbody>\n          </table>\n          <p>If your SEO is in good shape, you have a head start. But it's a head start, not the finish.</p>\n          <h2>Where SEO stops and agent readiness carries on</h2>\n          <h3>Being let in, not just crawled</h3>\n          <p>SEO teams check that Googlebot can crawl. They rarely check how bot protection treats ChatGPT-User, Perplexity-User or a browser agent. A site can rank well and still challenge every AI assistant at the door, because the rules that protect it from scrapers were tuned for a world where the only good bots were search engines. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a>.</p>\n          <h3>Content that exists without JavaScript</h3>\n          <p>Google renders JavaScript, so a page built entirely in the browser can still rank. Most AI crawlers and many assistants don't run JavaScript. They read the HTML your server sends, and if the price or description only appears after scripts load, they see an empty page. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see JavaScript-only content</a> and <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">prices AI agents can read</a>.</p>\n          <h3>A page that can be used, not just read</h3>\n          <p>SEO ends at the landing page. Agent readiness carries on through the page: buttons and links with names, form fields with labels, pop-ups that can be closed, and an add-to-cart button an agent can find and press. These look like accessibility concerns because they largely are. See <a href=\"https://ghostagentlab.com/blog/accessibility-is-agent-readiness/\">accessibility work is agent readiness work</a>.</p>\n          <h3>A task that can be finished</h3>\n          <p>The biggest difference is the last step. An agent sent to buy something either completes the purchase or it doesn't. No ranking report tells you which. The only way to know is to send a real agent through the journey and watch. That's what <a href=\"https://ghostagentlab.com/blog/introducing-ghost-agents/\">Ghost Agents</a> do, and it's why Task completion carries the most weight in <a href=\"https://ghostagentlab.com/blog/how-agentscore-works/\">AgentScore</a>.</p>\n          <h3>Measuring what agents bring</h3>\n          <p>SEO is measured in rankings, impressions and organic sessions. Agents mostly don't run your analytics tag, so they don't show up in those numbers. Agent visits live in server and CDN logs, and the visits and sales that AI assistants refer need their own tracking. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure agent traffic</a> and <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">tracking visits and sales that come from AI assistants</a>.</p>\n          <h2>Found versus chosen</h2>\n          <p>Here's the simplest way to put the difference to a leadership team.</p>\n          <ul>\n            <li><strong>SEO gets you found.</strong> It decides whether you're in the answer, and how prominently.</li>\n            <li><strong>Agent readiness gets you chosen.</strong> It decides whether an agent that found you can confirm the price, check the returns policy and complete the order, or whether it gives up and picks a site it can use.</li>\n          </ul>\n          <p>An assistant that can't read your stock level has no reason to recommend you over a competitor whose stock level is right there in the HTML. Being found and then failing the agent is close to not being found at all.</p>\n          <h2>What this means for SEO teams</h2>\n          <p>This isn't a new discipline that replaces SEO. It's SEO's natural next step, and SEO teams are well placed to lead it. They already own robots.txt, titles, sitemaps and much of the content. They already work with developers on structured data and rendering. What's new is a set of partners SEO hasn't always needed: whoever runs bot protection and the CDN, and whoever owns checkout. We cover that split in <a href=\"https://ghostagentlab.com/blog/who-owns-agent-readiness/\">who owns agent readiness?</a></p>\n          <div>\n            <p><strong>Three things an SEO lead can do this week:</strong></p>\n            <ol>\n              <li>Check robots.txt for rules that block AI assistants or AI search by accident, and decide separately what to do about training crawlers. Our <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">decision guide</a> helps.</li>\n              <li>View the source of a top product page and confirm the price, stock and description are in the HTML, not added later by JavaScript.</li>\n              <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> on your home page and share the Access findings with whoever runs your CDN or bot protection.</li>\n            </ol>\n          </div>\n          <p>If you want a structured path from there, our <a href=\"https://ghostagentlab.com/blog/90-day-agent-readiness-plan/\">90-day agent readiness plan</a> breaks the work into three phases.</p>",
      "date_published": "2026-10-09T12:22:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Perspective"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/how-agentscore-works/",
      "url": "https://ghostagentlab.com/blog/how-agentscore-works/",
      "title": "How AgentScore works: what's behind the 100 points",
      "summary": "How AgentScore turns its checks into a score out of 100: the four categories, how points are earned, which pages it visits, and what isn't scored yet.",
      "content_html": "<p>AgentScore gives your website one number out of 100 for how well AI agents can get in, read your pages and find their way around. A single number is only useful if you trust it, so here is exactly how it's worked out: the four categories, how each check earns points, which pages the scanner visits, and what it doesn't score yet.</p>\n          <h2>Four categories, 100 points</h2>\n          <p>Every check belongs to one of four categories. Each category has a fixed number of points, set by how much it matters to an AI agent trying to do something for a customer.</p>\n          <table>\n            <thead><tr><th>Category</th><th>Points</th><th>The question it answers</th></tr></thead>\n            <tbody>\n              <tr><td>Access</td><td>25</td><td>Can AI agents get in, past robots.txt, bot protection and challenges?</td></tr>\n              <tr><td>Readability</td><td>25</td><td>Can they understand your pages without guessing?</td></tr>\n              <tr><td>Navigability</td><td>20</td><td>Can they find and use the buttons, links and forms they need?</td></tr>\n              <tr><td>Task completion</td><td>30</td><td>Can an AI agent actually finish a real task, such as reaching checkout?</td></tr>\n            </tbody>\n          </table>\n          <p>The background to these categories is in <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">our guide to agent readiness</a>. This post is about the mechanics.</p>\n          <h2>How a check earns points</h2>\n          <p>Each check ends in one of four results:</p>\n          <ul>\n            <li><strong>Pass</strong> earns the check's full weight.</li>\n            <li><strong>Warning</strong> earns half. Something works, but agents will stumble on it.</li>\n            <li><strong>Fail</strong> earns nothing.</li>\n            <li><strong>Not tested</strong> is left out of the sum entirely. It never counts against you.</li>\n          </ul>\n          <p>Within a category, checks carry different weights, because some problems stop an agent cold and others only slow it down. A category's score is its points multiplied by the share of weight you earned among the checks that could be tested. If Readability's tested checks are worth 17 in total and you earn 13.5 of them, you get 25 &times; 13.5 &divide; 17, or 19.9 out of 25.</p>\n          <p>The overall score adds up the categories that could be tested and scales the result to 100. So if Task completion wasn't tested and you scored 20 for Access, 17 for Readability and 11 for Navigability, your AgentScore is 48 out of a possible 70, which is 69.</p>\n          <p>Most checks are rule-based rather than judged by an AI model, so the same site gets the same score from one scan to the next. When it moves, something on the site changed.</p>\n          <h2>What the bands mean</h2>\n          <table>\n            <thead><tr><th>Score</th><th>Band</th><th>What we tell you</th></tr></thead>\n            <tbody>\n              <tr><td>80 and up</td><td>Good</td><td>AI agents can use this site well.</td></tr>\n              <tr><td>60 to 79</td><td>Fair</td><td>AI agents can mostly use this site, but some things get in their way.</td></tr>\n              <tr><td>40 to 59</td><td>Weak</td><td>AI agents will struggle on this site.</td></tr>\n              <tr><td>Below 40</td><td>Poor</td><td>Most AI agents will fail on this site.</td></tr>\n            </tbody>\n          </table>\n          <p>The one-line verdict at the top of a report also names your biggest issue: the failed check with the highest weight.</p>\n          <h2>The checks, by category</h2>\n          <p>Here is what each category looks at today. The weight is how much the check counts inside its category.</p>\n          <h3>Access</h3>\n          <ul>\n            <li><strong>Bot protection lets AI agents through</strong> (weight 4). We request your home page as five AI agents and as a normal browser, and compare. If any agent is blocked or challenged while the browser gets through, it fails. If an agent gets less than half the browser's text, it's a warning.</li>\n            <li><strong>robots.txt lets AI assistants and search agents in</strong> (3). Blocking one or two assistants or AI search agents is a warning; three or more is a fail. Blocked training crawlers are reported but not scored.</li>\n            <li><strong>AI agents can open your product, pricing and cart pages</strong> (3).</li>\n            <li><strong>AI agents aren't blocked or shown a CAPTCHA at the cart and checkout</strong> (3).</li>\n            <li><strong>No CAPTCHA or challenge on arrival</strong> (2).</li>\n            <li><strong>Shoppers can check out without an account</strong> (2).</li>\n            <li><strong>llms.txt guide for AI</strong> (1), <strong>Agents can use your site through MCP or WebMCP</strong> (1) and <strong>Agents can check out through an agentic commerce protocol</strong> (1). These are newer standards, so a missing one is a warning, never a fail.</li>\n          </ul>\n          <h3>Readability</h3>\n          <ul>\n            <li><strong>Content loads without JavaScript</strong> (4): how much of the page's text is in the HTML before scripts run.</li>\n            <li><strong>Structured data describes your business and products</strong> (3) and <strong>Product pages give price and stock in a form agents can read</strong> (3).</li>\n            <li><strong>Clear page title, description, and headings</strong> (2), <strong>Valid sitemap</strong> (2) and <strong>Prices are in the page HTML</strong> (2).</li>\n            <li><strong>Images have text descriptions</strong> (1).</li>\n          </ul>\n          <h3>Navigability</h3>\n          <ul>\n            <li><strong>Buttons and links have names</strong> (4) and <strong>Pop-ups and banners can be dismissed by agents</strong> (4).</li>\n            <li><strong>Agents can find and press your add-to-cart or sign-up button</strong> (3) and <strong>Cart and checkout fields are labelled for agents</strong> (3).</li>\n            <li><strong>Form fields are labelled</strong> (2), <strong>Page has main and navigation landmarks</strong> (2), <strong>Clickable things are real buttons and links</strong> (2) and <strong>Key pages are linked from the home page</strong> (2).</li>\n          </ul>\n          <p>Weights can change between scoring versions as we learn what trips agents up. Every report records the scoring version it used, so you can compare like with like.</p>\n          <h2>What the scanner visits</h2>\n          <p>AgentScore reads public pages only. It never adds anything to a cart, submits a form, signs up or places an order. In one scan it looks at:</p>\n          <ol>\n            <li><strong>Your home page</strong>, three ways: as a browser sees it before JavaScript runs, rendered in a real browser, and as five AI agents (ChatGPT-User, Claude-User, Perplexity-User, OAI-SearchBot and Googlebot).</li>\n            <li><strong>Your robots.txt, sitemap and llms.txt</strong>, plus the well-known addresses where sites describe an MCP server or an agentic commerce profile.</li>\n            <li><strong>Up to three key pages</strong>: a product page, a pricing page and the cart. It looks for them the way an agent would, in the links on your home page first, then in your sitemap, then at common addresses such as <code>/pricing</code> and <code>/cart</code>. Each is fetched as a browser and as an AI assistant, and product and pricing pages are rendered in a browser too.</li>\n            <li><strong>Your checkout</strong>, found from the link on the cart page or at <code>/checkout</code>, fetched as a browser and as three AI assistants.</li>\n          </ol>\n          <p>Requests under the scanner's own name are signed so you can verify them. The requests that imitate AI agents are not, on purpose. We explain why in <a href=\"https://ghostagentlab.com/blog/scanner-wears-other-names/\">Why our scanner sometimes wears other agents' names</a>. How it identifies itself and how to opt out are on the <a href=\"https://ghostagentlab.com/agentscore/bot/\">AgentScore bot</a> page.</p>\n          <h2>What isn't scored yet</h2>\n          <div>\n            <p><strong>Task completion.</strong> The automated scan doesn't run an AI agent through a task yet, so \"An AI agent completes a real task\" shows as not tested and your score covers Access, Readability and Navigability. In the Ghost Agent Labs app, <a href=\"https://ghostagentlab.com/blog/introducing-ghost-agents/\">Ghost Agent tests</a> already send real AI agents through your journeys. Their results don't feed into AgentScore yet.</p>\n            <p><strong>Store-only checks on other sites.</strong> Guest checkout, checkout fields, product data and agentic commerce only apply to stores. If we find no products or cart, they're marked not tested.</p>\n            <p><strong>Anything behind a full cart.</strong> Many checkouts only open with items in the cart. Because the scanner never adds any, those checks often say \"not tested\" rather than guess.</p>\n            <p><strong>Your choice on training crawlers.</strong> Blocking AI training is a business decision, not a readiness problem, so it costs no points.</p>\n          </div>\n          <h2>How to use your score</h2>\n          <ol>\n            <li>Read the verdict and the biggest issue first. It's the highest-weight failure.</li>\n            <li>Fix failures before warnings, and Access before everything else: if agents can't get in, nothing else matters.</li>\n            <li>Hand each finding to the right person. Most Access problems belong to whoever runs your CDN or bot protection; most Navigability problems belong to developers.</li>\n            <li>Rescan after each fix. In the Ghost Agent Labs app you can rescan whenever you like. The free AgentScore page reuses a domain's result for 24 hours, so repeated requests don't hit your site again.</li>\n          </ol>\n          <p>Then go beyond the score: set up <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">tests of your key journeys with AI agents</a>, so you know when a release breaks them.</p>",
      "date_published": "2026-10-09T12:21:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Product"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/introducing-ghost-agents/",
      "url": "https://ghostagentlab.com/blog/introducing-ghost-agents/",
      "title": "Ghost Agents: testing your site with real AI agents",
      "summary": "Ghost Agents are real AI agents that run your checkout, search and sign-up journeys on a schedule, with pass rates, step-by-step replays and alerts.",
      "content_html": "<p>A readiness score tells you what might trip an AI agent up. It can't tell you whether one actually gets from your home page to checkout. Ghost Agents do. They're real AI agents that work through your key journeys the way a customer's assistant would, on a schedule, and show you every step.</p>\n          <h2>Why test with a real agent</h2>\n          <p>Checks for labelled buttons, structured data and bot protection catch the common problems. But journeys fail in ways no rule predicts: a size picker the agent can't operate, a discount pop-up that appears on the second page, a checkout button that only shows after scrolling. People get past these without noticing. Agents often don't.</p>\n          <p>This is what our team has done for human visitors for years with <a href=\"https://ghostinspector.com/\">Ghost Inspector</a>: prove the journeys that make you money still work, and find out when a release breaks one. Ghost Agents bring the same idea to AI agents.</p>\n          <h2>How a test works</h2>\n          <p>You write a goal in plain English, the way a customer might ask an AI assistant. To make it quick, there are ready-made templates for the journeys that matter most:</p>\n          <table>\n            <thead><tr><th>Template</th><th>What the agent tries to do</th></tr></thead>\n            <tbody>\n              <tr><td>Checkout</td><td>Find a popular product, add it to the cart, and check out as a guest up to the payment step</td></tr>\n              <tr><td>Product search</td><td>Use site search to find a product a customer might ask for, and open its page</td></tr>\n              <tr><td>Sign-up</td><td>Create an account with a test email address, as far as the site allows without email confirmation</td></tr>\n              <tr><td>Contact form</td><td>Find the contact or support form and fill it in with a short test question, without submitting it</td></tr>\n            </tbody>\n          </table>\n          <p>On each run, the agent opens your site in a fresh browser and repeats a simple loop. It looks at the page as an agent sees it: the address, the visible text, and the buttons, links and fields with their names. It picks one action, such as click, type, choose an option, scroll, go back or finish. Then it acts, and looks again.</p>\n          <h2>How you know it passed</h2>\n          <p>An agent saying \"done\" isn't proof. So you choose pass conditions that are checked against the final page the agent reaches, without asking an AI model:</p>\n          <ul>\n            <li>The page address contains some text, such as <code>/checkout</code>.</li>\n            <li>The page shows some text, such as \"Thanks for signing up\".</li>\n            <li>The agent reached the payment step. This is the one to use for checkout tests.</li>\n            <li>The agent reports the goal is done. This is used only if you set no other condition.</li>\n          </ul>\n          <p>AI agents don't behave the same way every time, so each check runs the journey several times (three is a good default) and reports a pass rate. A test where every run passes is passing. One where none pass is failing. One where results disagree is flaky: agents can do the journey, but not reliably, which for a customer's assistant can mean the same as not at all.</p>\n          <p>Tests run when you click \"Run now\", every day, or every hour. When one fails, you can be alerted by email or Slack.</p>\n          <h2>Replays: see exactly where it got stuck</h2>\n          <p>Every run is recorded. The replay opens on the step where things went wrong, with a screenshot of what the agent saw, the action it took, and its reasoning, which often says in plain words what it couldn't find or press. If a journey used to work, you can compare the failed run with the last passing one to see what changed.</p>\n          <p>That turns \"agents can't check out\" into something a developer can act on: \"the cookie banner covers the checkout button and has no labelled close button\".</p>\n          <h2>Safe by design</h2>\n          <p>A test agent that clicks around a live store has to be careful. Ghost Agents follow hard rules that the model can't talk its way past:</p>\n          <ul>\n            <li>They stop as soon as a payment form appears, and never enter card details. Reaching that point is how a checkout test passes.</li>\n            <li>They refuse buttons that place orders, pay, or delete or cancel an account.</li>\n            <li>They type only the test data you give the test, such as a test email address.</li>\n            <li>They stay on the site being tested.</li>\n            <li>Each run stops after 40 steps or five minutes, whichever comes first.</li>\n          </ul>\n          <p>They also identify themselves. Every request carries <code>GhostAgent/1.0 (+https://ghostagentlab.com/ghost-agent)</code> at the end of a normal Chrome user agent and is signed with Web Bot Auth, so your CDN can verify it's really us. We explain signing in <a href=\"https://ghostagentlab.com/blog/signed-requests/\">Why we sign every request our agents send</a>, and the <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> page covers how to allow or opt out.</p>\n          <h2>What Ghost Agents don't do yet</h2>\n          <p>We'd rather you knew the limits up front:</p>\n          <ul>\n            <li><strong>They don't feed AgentScore yet.</strong> Task completion is the AgentScore category these tests will fill. Until they're connected, the score covers Access, Readability and Navigability. See <a href=\"https://ghostagentlab.com/blog/how-agentscore-works/\">How AgentScore works</a>.</li>\n            <li><strong>One agent per run.</strong> Every agent gets the same instructions and tools, so results from different AI models can be compared, but running several side by side in one check isn't available yet.</li>\n            <li><strong>No release triggers yet.</strong> Tests run on demand or on a schedule, not automatically after each deploy.</li>\n            <li><strong>No payments.</strong> By design, a checkout test proves an agent can reach payment. It never completes one.</li>\n          </ul>\n          <h2>Where to start</h2>\n          <div>\n            <ol>\n              <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> and fix anything that keeps agents out, such as blocked assistants or a challenge page. A Ghost Agent can't test a journey it can't start.</li>\n              <li>In the app, add your site and create a Checkout test (or Sign-up, if you sell software). Run it once by hand.</li>\n              <li>Watch the replay, even if it passed. You'll learn how agents read your pages.</li>\n              <li>Set it to run daily and turn on alerts, so you hear about the release that breaks agent checkout before your customers' assistants do.</li>\n            </ol>\n          </div>\n          <p>For help choosing journeys and writing good goals, read <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">How to test your key journeys with AI agents</a>.</p>",
      "date_published": "2026-10-09T12:20:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Product"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/accessibility-is-agent-readiness/",
      "url": "https://ghostagentlab.com/blog/accessibility-is-agent-readiness/",
      "title": "Accessibility work is agent readiness work",
      "summary": "Many browser agents read pages through the accessibility tree. Where WCAG and agent readiness overlap, where they differ, and how to make fixes pay twice.",
      "content_html": "<p>If your team has spent time making your site work with screen readers and keyboards, you've already done a large share of the work to make it work for AI agents. The reason is simple: many browser agents read a page through the same structure that assistive technology uses.</p>\n          <h2>How a browser agent sees a page</h2>\n          <p>A person looks at a page and sees layout, color and images. A browser agent, the kind that opens your site in a real browser to search, fill in forms and add to cart, needs something it can reason about and act on. It has two main options.</p>\n          <ul>\n            <li><strong>Screenshots.</strong> The agent looks at an image of the page and decides where to click. This works, but it's slow, and it struggles with small icons, overlapping elements and anything that looks clickable but isn't.</li>\n            <li><strong>The accessibility tree.</strong> Browsers already build a structured version of every page for assistive technology: each heading, link, button and form field, with its role (what kind of thing it is), its name (what it's called) and its state (checked, expanded, disabled). Many agents read this tree, often alongside screenshots, because it tells them exactly what can be clicked and what each thing is for.</li>\n          </ul>\n          <p>The tooling used to build agents reflects this. Playwright's MCP server, a widely used way to give an AI model control of a browser, hands the model an accessibility snapshot of the page rather than pixels. In that snapshot, a button is a line like <code>button \"Add to cart\"</code>. A button with no name is just <code>button</code>, and the agent has to guess.</p>\n          <pre><code>&lt;!-- What the agent can use --&gt;\n&lt;button type=\"submit\"&gt;Add to cart&lt;/button&gt;\n&lt;button aria-label=\"Open cart\"&gt;&lt;svg aria-hidden=\"true\"&gt;…&lt;/svg&gt;&lt;/button&gt;\n&lt;!-- What it can't --&gt;\n&lt;div class=\"btn\" onclick=\"addToCart()\"&gt;&lt;svg&gt;…&lt;/svg&gt;&lt;/div&gt;</code></pre>\n          <p>The last example has no role and no name. A screen reader user can't use it, and an agent reading the accessibility tree may not know it exists.</p>\n          <h2>Where WCAG and agent readiness overlap</h2>\n          <p>The <a href=\"https://www.w3.org/TR/WCAG22/\">Web Content Accessibility Guidelines (WCAG) 2.2</a> are the standard most accessibility programs work to. Several of their success criteria line up closely with AgentScore checks.</p>\n          <table>\n            <thead><tr><th>WCAG 2.2 success criterion</th><th>Related AgentScore check</th></tr></thead>\n            <tbody>\n              <tr><td>4.1.2 Name, Role, Value</td><td>Buttons and links have names; Clickable things are real buttons and links</td></tr>\n              <tr><td>1.3.1 Info and Relationships, 3.3.2 Labels or Instructions</td><td>Form fields are labelled; Cart and checkout fields are labelled for agents</td></tr>\n              <tr><td>1.1.1 Non-text Content</td><td>Images have text descriptions</td></tr>\n              <tr><td>2.4.1 Bypass Blocks, 1.3.1 Info and Relationships</td><td>Page has main and navigation landmarks</td></tr>\n              <tr><td>2.4.2 Page Titled, 2.4.6 Headings and Labels</td><td>Clear page title, description, and headings</td></tr>\n              <tr><td>2.1.2 No Keyboard Trap, 2.4.11 Focus Not Obscured (Minimum)</td><td>Pop-ups and banners can be dismissed by agents</td></tr>\n            </tbody>\n          </table>\n          <p>The match isn't exact. WCAG asks more of each criterion than our checks test for, and our checks look at things from an agent's point of view. But the direction is the same: give every control a role and a name, label every field, describe every meaningful image, and structure the page so it can be navigated without seeing it.</p>\n          <p>Our guide to <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a> covers the fixes in detail, and <a href=\"https://ghostagentlab.com/articles/image-alt-text-ai-agents/\">alt text for AI agents</a> covers product images.</p>\n          <h2>Common accessibility gaps that also stop agents</h2>\n          <ul>\n            <li><strong>Icon-only buttons with no label.</strong> The cart, search and menu icons in a header are the usual culprits. Add visible text or an <code>aria-label</code>.</li>\n            <li><strong>Placeholder text used as a label.</strong> It disappears when the field is filled in and isn't a reliable name. Use a real <code>&lt;label&gt;</code>.</li>\n            <li><strong>Clickable boxes.</strong> A <code>div</code> with a click handler looks like a button to a person and like plain text to the accessibility tree. Use <code>&lt;button&gt;</code> or <code>&lt;a href&gt;</code>.</li>\n            <li><strong>Menus that only open on hover.</strong> Neither keyboard users nor many agents can hover. Open them on click and focus too.</li>\n            <li><strong>Custom size and color pickers.</strong> If a swatch has no name and no selected state, an agent can't tell which size it chose. Native radio buttons or properly labelled custom controls fix this.</li>\n            <li><strong>Pop-ups without a labelled close button.</strong> See <a href=\"https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/\">cookie banners and pop-ups that trap AI agents</a>.</li>\n          </ul>\n          <h2>Where they differ</h2>\n          <p>Accessibility work won't cover everything, and some of it doesn't matter to agents at all.</p>\n          <p><strong>Agent readiness needs more than accessibility.</strong> A perfectly accessible site can still block AI assistants in robots.txt, challenge them with bot protection, or load its prices with JavaScript that many AI crawlers never run. Screen readers work inside a real browser, so they see JavaScript content; many agents don't. Structured data, sitemaps and llms.txt matter to agents and have little to do with accessibility. That's why AgentScore has Access and Readability categories as well as Navigability.</p>\n          <p><strong>Some accessibility work matters less to agents.</strong> Color contrast, text resizing, captions and motion settings are essential for people and largely irrelevant to software. Don't use agent readiness as a reason to deprioritize them. They serve customers, and in many places they're a legal requirement.</p>\n          <p><strong>Bad ARIA hurts both.</strong> The W3C's own guidance is that no ARIA is better than bad ARIA. An <code>aria-label</code> that says \"button\", a <code>role=\"button\"</code> on something that can't be pressed, or <code>aria-hidden=\"true\"</code> on a real control misleads screen readers and agents alike. Native HTML elements are almost always the better choice.</p>\n          <div>\n            <p><strong>A note for leaders:</strong> in the EU, the European Accessibility Act has applied to many e-commerce services since June 2025, so many retailers already have accessibility programs underway. If yours does, put agent readiness alongside it rather than starting a separate project. The same fixes, the same developers and often the same tickets.</p>\n          </div>\n          <h2>How to use this</h2>\n          <ol>\n            <li><strong>Ask your accessibility lead for the latest audit.</strong> Unlabelled controls, missing form labels and keyboard traps in that report are agent readiness issues too. Fixing them pays twice.</li>\n            <li><strong>Add an agent's view to your testing.</strong> Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> and compare its Navigability findings with your audit. Then send a <a href=\"https://ghostagentlab.com/blog/introducing-ghost-agents/\">Ghost Agent</a> through checkout to see whether a real agent can finish.</li>\n            <li><strong>Make it part of the definition of done.</strong> New components ship with names, labels and native elements. That's cheaper than fixing them later, for people and agents alike.</li>\n          </ol>\n          <p>Accessibility and agent readiness come from the same idea: a website should work for visitors who don't see it the way its designers do. More and more of those visitors are software acting for a person.</p>",
      "date_published": "2026-10-09T12:19:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Perspective"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/scanner-wears-other-names/",
      "url": "https://ghostagentlab.com/blog/scanner-wears-other-names/",
      "title": "Why our scanner sometimes wears other agents' names",
      "summary": "AgentScore requests a few pages as ChatGPT, Claude, Perplexity and Googlebot, unsigned, to see how your site treats them. Here's exactly what and why.",
      "content_html": "<p>If you look closely at your logs during an AgentScore scan, you'll see a few requests that say they're ChatGPT, Claude, Perplexity or Googlebot. They aren't. They're our scanner, borrowing those agents' user agents to see how your site treats them. Here's why we do it, exactly what we send, and what the result can and can't tell you.</p>\n          <h2>The question we're trying to answer</h2>\n          <p>One of the most common reasons an AI agent fails on a website has nothing to do with the website's design. The agent is turned away at the door. A bot protection rule, a firewall setting or a CDN feature decides that anything calling itself an AI agent gets a challenge page, an error, or a stripped-down page with no prices.</p>\n          <p>Nobody notices, because people never see it. The site works perfectly in a browser. The only way to find out is to knock as the agent and compare what comes back with what a browser gets.</p>\n          <p>We can't send the real ChatGPT to your site on demand. So for a handful of requests, our scanner sends the same <code>User-Agent</code> header those agents send, and compares the answer with a normal browser's.</p>\n          <h2>Exactly what we send</h2>\n          <table>\n            <thead><tr><th>Name we use</th><th>What we request</th><th>Signed?</th></tr></thead>\n            <tbody>\n              <tr><td><code>AgentScore/1.0</code>, our own name</td><td>robots.txt, sitemaps, llms.txt, well-known files (MCP and agentic commerce discovery), and checks for common pages such as <code>/pricing</code> and <code>/cart</code></td><td>Yes</td></tr>\n              <tr><td>A normal Chrome browser</td><td>The home page, your llms.txt, and the key pages, before and after JavaScript runs. This is the baseline everything is compared with</td><td>No</td></tr>\n              <tr><td>ChatGPT-User, Claude-User, Perplexity-User, OAI-SearchBot and Googlebot</td><td>The home page, once each</td><td>No</td></tr>\n              <tr><td>ChatGPT-User</td><td>The product, pricing and cart pages we found, once each</td><td>No</td></tr>\n              <tr><td>ChatGPT-User, Claude-User and Perplexity-User</td><td>The cart and checkout, once each</td><td>No</td></tr>\n            </tbody>\n          </table>\n          <p>The imitated requests use the user agent strings those operators publish, so they look the way the real thing does. We picked these agents because they're the ones that fetch pages for people in real time or power AI search, which are the visits most likely to turn into customers.</p>\n          <p>Training crawlers such as GPTBot and ClaudeBot are handled differently. We read your robots.txt rules for them, along with every other AI agent we track, but we don't request pages as them. Blocking training is a business choice, and it doesn't cost points. Our guide to <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">blocking or allowing AI crawlers</a> covers that decision.</p>\n          <h2>How we compare the answers</h2>\n          <p>For each imitated request, we line the response up against the browser's:</p>\n          <ul>\n            <li><strong>Blocked:</strong> the agent got an error (HTTP 400 or above) where the browser didn't, or a challenge page such as \"Just a moment\" or \"Verify you are human\" where the browser got the real page.</li>\n            <li><strong>Degraded:</strong> the agent got less than half the visible text the browser got, which usually means a placeholder or stripped-down page.</li>\n            <li><strong>Same page:</strong> neither of the above.</li>\n          </ul>\n          <p>On the home page, any blocked agent fails the <strong>Bot protection lets AI agents through</strong> check, and a degraded one is a warning. A similar comparison, looking for errors and challenge pages (and, at checkout, CAPTCHAs), runs on your product, pricing and cart pages for <strong>AI agents can open your product, pricing and cart pages</strong>, and on the cart and checkout for <strong>AI agents aren't blocked or shown a CAPTCHA at the cart and checkout</strong>. <a href=\"https://ghostagentlab.com/blog/how-agentscore-works/\">How AgentScore works</a> explains how those checks add up.</p>\n          <h2>Why those requests aren't signed</h2>\n          <p>Everything the scanner sends under its own name is signed with Web Bot Auth, so your CDN can prove it came from us. We explain how in <a href=\"https://ghostagentlab.com/blog/signed-requests/\">Why we sign every request our agents send</a>.</p>\n          <p>The imitated requests are deliberately unsigned. If we signed them, a CDN that recognizes and trusts our signature might wave them through, and we'd be measuring how your site treats Ghost Agent Labs, not how it treats ChatGPT. Unsigned, they get exactly the treatment any request with that name gets.</p>\n          <h2>What the result can't tell you</h2>\n          <p>Our imitation is honest about one thing it can't do: it can't pass identity checks. The requests come from our servers, not from OpenAI's, Anthropic's, Perplexity's or Google's, and they carry none of those companies' signatures.</p>\n          <p>So if your bot protection verifies agents properly, checking published IP ranges or reverse DNS and blocking impostors, it may block our imitation too. That's the right behavior, and it will show up as a failed check. When that happens:</p>\n          <ol>\n            <li>Look at the evidence in the report. It shows which agents were blocked and what status they got.</li>\n            <li>Check your CDN or bot protection settings: is the rule \"block requests that claim to be an AI agent but fail verification\", or \"block AI agents\"? The first is good practice. The second turns customers away.</li>\n            <li>Confirm in your logs, or on the verification page in the Ghost Agent Labs app, that verified requests from the real agent are getting through.</li>\n          </ol>\n          <p>If the real agents get through and only impostors are blocked, you're in good shape, whatever the check says. Our guides to <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">telling whether an AI crawler is real</a> and <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection that lets AI agents through</a> cover how to set that up.</p>\n          <p>The opposite limit applies too. A site that treats these agents well when they come from our servers will probably treat the real ones well, but real products differ in details we can't copy, such as where their requests come from. Treat the result as a strong signal, not a guarantee.</p>\n          <h2>Keeping it small</h2>\n          <p>Imitated requests are a small part of a scan: a few page loads per agent at most, never a crawl. The scanner reads public pages only and never adds to a cart, submits a form or places an order. On the free AgentScore page, a domain's result is reused for 24 hours, so repeated requests for the same site don't cause repeated visits.</p>\n          <div>\n            <p><strong>If you see these requests in your logs:</strong> they'll arrive close together with a scan from <code>AgentScore/1.0</code>. The <a href=\"https://ghostagentlab.com/agentscore/bot/\">AgentScore bot</a> page explains how to verify our signed requests and how to opt out, including from the comparison requests.</p>\n            <p><strong>If you want to see the result:</strong> run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> on your own site and open the Access checks.</p>\n          </div>",
      "date_published": "2026-10-09T12:18:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Engineering"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/",
      "url": "https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/",
      "title": "Should you block AI crawlers? A decision guide for leaders",
      "summary": "Blocking AI crawlers is several decisions, not one. How to decide on training crawlers, AI search and assistants, and how to enforce the policy.",
      "content_html": "<p>\"Should we block AI?\" usually reaches a leadership meeting as one question. It's really three or four, because \"AI crawlers\" covers several kinds of software doing very different jobs. Block the wrong one and you vanish from the answers your customers are reading. Here's how to make the call, kind by kind.</p>\n          <h2>Start by splitting the question</h2>\n          <p>The major AI companies each run several agents, and each has its own name so you can treat them differently.</p>\n          <table>\n            <thead><tr><th>Kind</th><th>What it does</th><th>Examples</th></tr></thead>\n            <tbody>\n              <tr><td>Training crawlers</td><td>Collect pages to train future AI models</td><td>GPTBot, ClaudeBot, CCBot, Google-Extended (a control name, not a separate crawler)</td></tr>\n              <tr><td>AI search crawlers</td><td>Index pages so AI search can cite and link to you</td><td>OAI-SearchBot, Claude-SearchBot, PerplexityBot</td></tr>\n              <tr><td>Assistant fetchers</td><td>Fetch a page in real time because a person asked about it</td><td>ChatGPT-User, Claude-User, Perplexity-User</td></tr>\n              <tr><td>Browser agents</td><td>Drive a real browser for a person: search, fill in forms, add to cart</td><td>Usually look like an ordinary browser</td></tr>\n            </tbody>\n          </table>\n          <p>Our <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt guide</a> has the longer list and the exact names. The point for decision-makers is simpler: you can say yes to some and no to others.</p>\n          <h2>The case for each, in business terms</h2>\n          <h3>Assistant fetchers: almost always allow</h3>\n          <p>Each visit stands in for a real person asking about you right now: \"Is this in stock?\", \"What's their returns policy?\", \"Which of these three is cheapest?\". Block them and the assistant answers without you, or recommends a competitor. For a store or a software company, this is the closest thing to a customer walking in.</p>\n          <h3>AI search crawlers: allow, unless you also opt out of search</h3>\n          <p>These build the index AI search draws on. They're the AI equivalent of being in Google's index, and they're how you get cited and linked. Few businesses that want search traffic would block Googlebot. The same logic applies here.</p>\n          <h3>Training crawlers: a real choice</h3>\n          <p>This is the one worth debating. Allowing training means your content may shape what AI models know, including about your brand and products. Blocking it keeps your content out of future training sets, and leaves AI search and assistants unaffected, because they use different names.</p>\n          <ul>\n            <li><strong>Lean towards blocking</strong> if your content is the product: publishers, research, courses, original reviews, anything you license or sell.</li>\n            <li><strong>Lean towards allowing</strong> if your content mainly exists to sell something else: product pages, help centers, pricing. Being well understood by AI models is usually worth more to you than the content itself.</li>\n          </ul>\n          <p>Either answer is legitimate. That's why AgentScore reports blocked training crawlers but doesn't take points off for them.</p>\n          <h3>Browser agents: you can't block them by name, so make them work</h3>\n          <p>Browser agents mostly look like a normal browser, so robots.txt rules by name don't reach them. The useful question isn't whether to block them but whether they can complete a purchase or sign-up when they arrive. That's what <a href=\"https://ghostagentlab.com/blog/introducing-ghost-agents/\">Ghost Agent tests</a> check.</p>\n          <h2>A quick decision guide</h2>\n          <table>\n            <thead><tr><th>If you are...</th><th>Assistants and AI search</th><th>Training crawlers</th></tr></thead>\n            <tbody>\n              <tr><td>An online store</td><td>Allow</td><td>Usually allow; your product pages are there to be read</td></tr>\n              <tr><td>A software or services company</td><td>Allow</td><td>Usually allow marketing and docs; your call on anything gated</td></tr>\n              <tr><td>A publisher or content business</td><td>Allow if you want AI citations and referral traffic</td><td>Often block, or allow only under a licensing agreement</td></tr>\n              <tr><td>Unsure</td><td>Allow</td><td>Decide deliberately; don't inherit someone's copied list</td></tr>\n            </tbody>\n          </table>\n          <h2>Three mistakes that cost more than the decision</h2>\n          <ol>\n            <li><strong>Copying a \"block all AI\" list.</strong> These lists usually include assistant fetchers and AI search crawlers, so you drop out of AI answers along with training.</li>\n            <li><strong>Deciding in robots.txt and forgetting the CDN.</strong> Bot protection, firewall rules and one-click \"block AI bots\" settings can override a friendly robots.txt. The agent never gets far enough to read it. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection that lets AI agents through</a>.</li>\n            <li><strong>Trusting the name.</strong> Anyone can claim to be ChatGPT or Googlebot. robots.txt is a request, not a lock: well-behaved agents follow it and scrapers ignore it. Stopping abusive traffic, including impostors using trusted names, has to happen at your CDN, by verifying who's really asking. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</li>\n          </ol>\n          <p>One more nuance: some operators treat an assistant fetch, made because a person asked about a specific page, like a person clicking a link, and say robots.txt may not apply to it. If you truly need to stop those visits, that's a CDN rule, not a robots.txt line. Most businesses want those visits most of all.</p>\n          <h2>What to do this month</h2>\n          <div>\n            <ol>\n              <li>Agree a written policy for each kind of agent: one line each, signed off by marketing, e-commerce and whoever owns your content rights.</li>\n              <li>Have someone write it into robots.txt. Allowing assistants and search while opting out of training looks like this:\n                <pre><code>User-agent: GPTBot\nUser-agent: ClaudeBot\nUser-agent: CCBot\nUser-agent: Google-Extended\nDisallow: /\nUser-agent: *\nAllow: /\nSitemap: https://northwind.example/sitemap.xml</code></pre>\n              </li>\n              <li>Ask whoever runs your CDN or bot protection to confirm its settings match the policy, and that it checks agents are genuine rather than trusting their names.</li>\n              <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>. It reads your robots.txt as the AI agents we track would, and requests your pages as several AI assistants to see whether your bot protection lets them through. <a href=\"https://ghostagentlab.com/blog/scanner-wears-other-names/\">Here's how that works</a>.</li>\n              <li>Measure what happens next. <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">Count the agents that visit</a> and <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">the visits and sales AI assistants send you</a>, and revisit the policy every quarter.</li>\n            </ol>\n          </div>\n          <p>The default for most businesses is simple: let the agents that bring customers in, make a deliberate choice about training, and enforce both where it actually counts.</p>",
      "date_published": "2026-10-09T12:17:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Perspective"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/90-day-agent-readiness-plan/",
      "url": "https://ghostagentlab.com/blog/90-day-agent-readiness-plan/",
      "title": "A 90-day agent readiness plan for e-commerce teams",
      "summary": "A practical three-phase plan: let AI agents in, make pages readable and usable, then prove agents can finish checkout, mapped to AgentScore checks.",
      "content_html": "<p>Agent readiness can feel like a long list of unrelated fixes. It's easier to run as three phases of about a month each: let AI agents in, help them understand and move around your site, then prove they can finish the jobs that make you money. Here's a plan an e-commerce team can start on Monday.</p>\n          <h2>Before you start</h2>\n          <p>Name one person to own the plan. It doesn't have to be a developer; on most teams it's someone in e-commerce, digital or SEO who can get time from the people who make the changes. Our post on <a href=\"https://ghostagentlab.com/blog/who-owns-agent-readiness/\">who owns agent readiness</a> shows who usually fixes what.</p>\n          <p>Then agree what \"done\" means. We suggest three outcomes by day 90:</p>\n          <ul>\n            <li>No critical AgentScore findings on your home page and key product pages.</li>\n            <li>At least one money journey, such as product to checkout, tested by a real AI agent on a schedule.</li>\n            <li>A monthly view of agent traffic and AI referrals that leadership can read.</li>\n          </ul>\n          <p>Notice that none of these is a target score. Scores are useful for tracking progress, but a site that loses ten points to an optional new standard can still serve agents well, and a site that scores well can still fail at the size selector. Aim at the outcomes.</p>\n          <h2>Days 1–30: let agents in and get a baseline</h2>\n          <p>The first month is about Access, because nothing else matters if agents are turned away at the door. It's also when you set the baseline you'll measure against.</p>\n          <h3>Week 1: measure</h3>\n          <ol>\n            <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> on your home page and save the report. <a href=\"https://ghostagentlab.com/blog/how-agentscore-works/\">How AgentScore works</a> explains what's behind each category.</li>\n            <li>Connect a data source, such as your CDN or server logs, so you can see which AI agents visit and how your site answers them. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure agent traffic</a>.</li>\n            <li>List your three most valuable journeys: for most stores, search to product, product to cart, and cart to checkout.</li>\n          </ol>\n          <h3>Weeks 2–4: fix Access</h3>\n          <table>\n            <thead><tr><th>AgentScore check</th><th>Usually fixed by</th><th>Read</th></tr></thead>\n            <tbody>\n              <tr><td>robots.txt lets AI assistants and search agents in</td><td>SEO or content team</td><td><a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a></td></tr>\n              <tr><td>Bot protection lets AI agents through</td><td>Bot protection or CDN admin</td><td><a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">Tell if an AI crawler is real</a></td></tr>\n              <tr><td>No CAPTCHA or challenge on arrival</td><td>Bot protection or CDN admin</td><td><a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">Bot protection and CAPTCHAs</a></td></tr>\n              <tr><td>AI agents can open your product, pricing and cart pages</td><td>Bot protection or CDN admin</td><td><a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">Rate limits for AI agents</a></td></tr>\n              <tr><td>llms.txt guide for AI</td><td>SEO or content team</td><td><a href=\"https://ghostagentlab.com/articles/llms-txt/\">How to write an llms.txt file</a></td></tr>\n            </tbody>\n          </table>\n          <p>Make one policy decision this month too: what you'll do about AI training crawlers, as distinct from the assistants that fetch pages for shoppers. Our <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">decision guide</a> walks through it. Settle it early, because it affects robots.txt and bot protection rules at the same time.</p>\n          <div>\n            <p><strong>End of month one:</strong> rescan. The AI assistants you want can reach your key pages, you've decided your position on training crawlers, and you have a baseline score and a first look at agent traffic.</p>\n          </div>\n          <h2>Days 31–60: make pages readable and usable</h2>\n          <p>Month two covers Readability and Navigability: whether agents can understand your pages, and whether they can find their way around them. Most of this is developer work, so get it into the sprint plan early in the month.</p>\n          <h3>Readability</h3>\n          <ul>\n            <li><strong>Content loads without JavaScript</strong> and <strong>Prices are in the page HTML.</strong> The highest-value fixes for most stores. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">JavaScript-only content</a> and <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">prices AI agents can read</a>.</li>\n            <li><strong>Structured data</strong> and <strong>product pages give price and stock in a form agents can read.</strong> See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data AI shopping agents can read</a>.</li>\n            <li><strong>Clear page title, description, and headings</strong>, <strong>image text descriptions</strong> and a <strong>valid sitemap.</strong> Usually SEO and content work. See <a href=\"https://ghostagentlab.com/articles/titles-descriptions-headings-ai/\">titles, descriptions and headings</a>, <a href=\"https://ghostagentlab.com/articles/image-alt-text-ai-agents/\">alt text for AI agents</a> and <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">XML sitemaps</a>.</li>\n            <li><strong>Policy pages</strong> for shipping and returns that an assistant can quote accurately. See <a href=\"https://ghostagentlab.com/articles/policy-pages-ai-assistants/\">shipping, returns and FAQ pages</a>.</li>\n          </ul>\n          <h3>Navigability</h3>\n          <ul>\n            <li><strong>Buttons and links have names</strong>, <strong>form fields are labelled</strong> and <strong>clickable things are real buttons and links.</strong> See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>. If you have an accessibility audit, start there: <a href=\"https://ghostagentlab.com/blog/accessibility-is-agent-readiness/\">much of it overlaps</a>.</li>\n            <li><strong>Pop-ups and banners can be dismissed by agents.</strong> See <a href=\"https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/\">cookie banners and pop-ups</a>.</li>\n            <li><strong>Main and navigation landmarks</strong> and <strong>key pages linked from the home page.</strong> See <a href=\"https://ghostagentlab.com/articles/link-key-pages-home-page/\">link your key pages from the home page</a>.</li>\n            <li><strong>Site search and filters</strong> that agents can operate. See <a href=\"https://ghostagentlab.com/articles/site-search-filters-ai-agents/\">site search and filters AI agents can use</a>.</li>\n          </ul>\n          <div>\n            <p><strong>End of month two:</strong> rescan and compare category by category with your baseline. Agents should be able to read the price, stock and description on your product pages from the HTML alone, and move around without getting stuck on unnamed controls or pop-ups.</p>\n          </div>\n          <h2>Days 61–90: prove agents can finish the job</h2>\n          <p>Month three is about Task completion, the category that carries the most weight. Checks can tell you a page looks usable. Only a real agent can tell you whether it is.</p>\n          <ol>\n            <li><strong>Test your money journeys.</strong> Set up <a href=\"https://ghostagentlab.com/blog/introducing-ghost-agents/\">Ghost Agent</a> tests for the journeys you listed in week one, starting with product to checkout. Each check runs the agent several times, because agents vary, and Ghost Agents stop before paying. See <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">how to test your key journeys with AI agents</a>.</li>\n            <li><strong>Fix where they fail.</strong> The usual sticking points are the add-to-cart button, option selectors, checkout fields and a challenge at the cart. The relevant checks are <strong>Agents can find and press your add-to-cart or sign-up button</strong>, <strong>Cart and checkout fields are labelled for agents</strong>, and <strong>AI agents aren't blocked or shown a CAPTCHA at the cart and checkout.</strong> See <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">agent-ready checkout</a>.</li>\n            <li><strong>Review guest checkout.</strong> An agent buying for someone can't easily create an account for them. See <a href=\"https://ghostagentlab.com/articles/guest-checkout-ai-agents/\">guest checkout: why AI shopping agents need it</a>.</li>\n            <li><strong>Look ahead.</strong> Ask your platform and developers where you stand on <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP and WebMCP</a> and on <a href=\"https://ghostagentlab.com/articles/agentic-commerce-protocols/\">agentic commerce protocols</a>. These are early and still changing, so they're weighted lightly in AgentScore. A decision and a plan are enough for now.</li>\n            <li><strong>Set up reporting.</strong> Track the visits and sales AI assistants send you, alongside agent traffic from your logs. See <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">tracking visits and sales that come from AI assistants</a>.</li>\n          </ol>\n          <div>\n            <p><strong>End of month three:</strong> at least one money journey runs on a schedule with alerts, failures go to a named owner, and leadership gets a one-page monthly view: AgentScore by category, Ghost Agent pass rates, agent traffic and AI referrals.</p>\n          </div>\n          <h2>After day 90</h2>\n          <p>Agent readiness drifts. A new pop-up campaign, a bot protection rule change or a redesigned product page can undo months of work in one release. Keep three habits:</p>\n          <ul>\n            <li>Rescan after every significant release, and at least monthly.</li>\n            <li>Keep Ghost Agent tests running on your money journeys, and add one when you launch a new one.</li>\n            <li>Add agent checks to your definition of done: named controls, labelled fields, prices in the HTML, and no new challenge pages on key paths.</li>\n          </ul>\n          <p>Start with a scan. It takes under a minute, and it'll tell you which month of this plan needs the most attention.</p>",
      "date_published": "2026-10-09T12:16:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Guide"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/who-owns-agent-readiness/",
      "url": "https://ghostagentlab.com/blog/who-owns-agent-readiness/",
      "title": "Who owns agent readiness?",
      "summary": "Agent readiness spans e-commerce, SEO, developers, security and legal. Who usually fixes what, and how to give the work one clear owner.",
      "content_html": "<p>Ask who owns agent readiness and you'll often get a pause. SEO thinks it's a developer problem. Developers think it's a security setting. Security thinks it's marketing's call. Meanwhile AI agents are turned away, or get lost, and nobody hears about it. The work is spread across several teams by its nature. What it needs is one clear owner and a clear split of the rest.</p>\n          <h2>Why it falls between teams</h2>\n          <p>Agent readiness touches four layers of a website, and each layer has its own team.</p>\n          <ul>\n            <li><strong>The door:</strong> robots.txt, bot protection, CDN rules and rate limits decide whether an agent gets in.</li>\n            <li><strong>The content:</strong> titles, descriptions, product data and policies decide whether it understands what it finds.</li>\n            <li><strong>The page:</strong> buttons, forms, pop-ups and rendering decide whether it can move around and act.</li>\n            <li><strong>The transaction:</strong> cart, checkout and accounts decide whether it can finish the job.</li>\n          </ul>\n          <p>No one team sees all four. Each sees its own layer working, and the failure shows up somewhere else: a lost recommendation, or an order that never happened.</p>\n          <h2>Who usually fixes what</h2>\n          <p>In the Ghost Agent Labs app, each AgentScore finding names a suggested owner under “Usually fixed by”, and the CSV export includes it too. Grouped together, the defaults look like this.</p>\n          <table>\n            <thead><tr><th>Team</th><th>What they usually own</th><th>AgentScore checks</th></tr></thead>\n            <tbody>\n              <tr><td>SEO or content team</td><td>The rules and text agents read first</td><td>robots.txt lets AI assistants and search agents in; llms.txt guide for AI; Clear page title, description, and headings; Images have text descriptions; Valid sitemap</td></tr>\n              <tr><td>Bot protection or CDN admin</td><td>Who gets through the door, and where</td><td>Bot protection lets AI agents through; No CAPTCHA or challenge on arrival; AI agents can open your product, pricing and cart pages; AI agents aren't blocked or shown a CAPTCHA at the cart and checkout</td></tr>\n              <tr><td>Developer</td><td>How pages are built and rendered</td><td>Content loads without JavaScript; Structured data describes your business and products; Product pages give price and stock in a form agents can read; Prices are in the page HTML; every Navigability check, from named buttons and labelled fields to pop-ups and checkout fields; Agents can use your site through MCP or WebMCP</td></tr>\n              <tr><td>E-commerce platform admin</td><td>How the store and checkout are configured</td><td>Shoppers can check out without an account; Agents can check out through an agentic commerce protocol</td></tr>\n            </tbody>\n          </table>\n          <p>Task completion, which comes from real agents running journeys, usually lands with developers once a test shows where an agent gets stuck. But the decision about which journeys matter belongs to the business.</p>\n          <p>These are defaults, not rules. On a hosted platform, the \"developer\" fixes for a product page may really be theme settings an e-commerce manager can change. At a small company, one person may hold three of these roles.</p>\n          <h2>The roles in more detail</h2>\n          <h3>E-commerce and digital leaders</h3>\n          <p>They own the outcome: whether AI agents can find, choose and buy from the site. That makes them the natural overall owner. They decide which journeys matter, set priorities when fixes compete for developer time, and own the platform settings that shape checkout, such as guest checkout. See <a href=\"https://ghostagentlab.com/articles/guest-checkout-ai-agents/\">guest checkout: why AI shopping agents need it</a>.</p>\n          <h3>Marketing and SEO</h3>\n          <p>They own most of what agents read before they act: robots.txt, titles and headings, alt text, sitemaps and llms.txt. They also own measurement, because AI referrals belong next to search and social in channel reporting. See <a href=\"https://ghostagentlab.com/blog/seo-to-agent-readiness/\">SEO got you found. Agent readiness gets you chosen</a> and <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">tracking visits and sales that come from AI assistants</a>.</p>\n          <h3>Developers</h3>\n          <p>They own the largest number of fixes: rendering content and prices in the HTML, structured data, named controls, labelled fields, dismissible pop-ups and checkout fields. Much of this overlaps with accessibility, so the same people and practices often apply. See <a href=\"https://ghostagentlab.com/blog/accessibility-is-agent-readiness/\">accessibility work is agent readiness work</a>. Developers also run <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">agent tests</a> as part of release checks, and evaluate newer interfaces such as <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP and WebMCP</a>.</p>\n          <h3>Security and CDN</h3>\n          <p>They own the door. Bot protection, WAF rules, challenge pages and rate limits are where agents are most often blocked by accident, because those rules were written to stop scrapers. This team's job isn't to let everything in. It's to tell real agents from impostors, by verified identity rather than by name alone, and to treat the agents the business wants accordingly. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a>, <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a> and <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a>.</p>\n          <h3>Legal</h3>\n          <p>Legal doesn't fix findings, but it shapes several decisions the other teams can't make alone:</p>\n          <ul>\n            <li>Whether to allow AI training crawlers, which is a content-licensing question as much as a technical one. See <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">should you block AI crawlers?</a></li>\n            <li>What your terms of service say about automated access, and whether they accidentally forbid the agents you want.</li>\n            <li>The terms and liability around purchases made by an agent for a customer, before you adopt an <a href=\"https://ghostagentlab.com/articles/agentic-commerce-protocols/\">agentic commerce protocol</a>.</li>\n            <li>Accessibility obligations, which often share fixes with agent readiness.</li>\n          </ul>\n          <h2>Making it work in practice</h2>\n          <ol>\n            <li><strong>Name one accountable owner.</strong> Usually the head of e-commerce or digital. They don't make every fix; they make sure each one has someone.</li>\n            <li><strong>Route findings by owner.</strong> Export the findings as CSV and send each group to its team. Each finding explains why it matters in plain language, which helps when the team receiving it doesn't think of AI agents as its problem.</li>\n            <li><strong>Agree the policy decisions once.</strong> Training crawlers, verified agents through bot protection, and guest checkout are decisions, not tickets. Settle them in one meeting with SEO, security, e-commerce and legal in the room.</li>\n            <li><strong>Make failures visible to the owner.</strong> Run <a href=\"https://ghostagentlab.com/blog/introducing-ghost-agents/\">Ghost Agent</a> tests on your money journeys and send alerts to the person who can act, not a shared inbox.</li>\n            <li><strong>Review monthly.</strong> AgentScore by category, Ghost Agent pass rates, agent traffic and AI referrals, on one page.</li>\n          </ol>\n          <div>\n            <p><strong>A simple test:</strong> if an AI assistant started being blocked from your product pages tomorrow, who would find out, and how? If the answer is \"nobody\" or \"eventually\", that's the gap to close first.</p>\n          </div>\n          <p>For a phased version of this, see our <a href=\"https://ghostagentlab.com/blog/90-day-agent-readiness-plan/\">90-day agent readiness plan</a>. Or start by running <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> and seeing whose name comes up most.</p>",
      "date_published": "2026-10-09T12:15:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Perspective"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/mcp-webmcp-for-websites/",
      "url": "https://ghostagentlab.com/articles/mcp-webmcp-for-websites/",
      "title": "MCP and WebMCP: giving AI agents a direct way into your site",
      "summary": "What MCP servers and WebMCP tools are, what they mean for your business, how AgentScore detects them, and how to start safely.",
      "content_html": "<p>Most AI agents use your website the way a person does: they read pages, click links and fill in forms. That works, but it's slow and it breaks easily. MCP and WebMCP give agents a second way in, a set of named actions such as \"search products\" or \"check order status\" that they can call directly. This guide explains both, what they mean for your business, and how to start small.</p>\n          <h2>Two ways an agent can use your site</h2>\n          <p>Today, an AI agent asked to find a waterproof jacket in a medium on your store has to work your pages. It loads the home page, finds the search box, reads the results, opens a product, works out the size picker, and reads the price. Every step depends on your page layout. A redesigned menu or a new pop-up can stop it.</p>\n          <p>The alternative is to tell the agent what it can do and let it ask directly. Instead of driving the search box, it calls a <code>search_products</code> action with \"waterproof jacket, medium\" and gets back a clean list of products, prices and stock. That's the idea behind both MCP and WebMCP.</p>\n          <ul>\n            <li><strong>Faster.</strong> One request instead of a dozen page loads and clicks.</li>\n            <li><strong>More reliable.</strong> Actions don't change when you redesign the page.</li>\n            <li><strong>More accurate.</strong> You decide exactly what data comes back, so agents quote the right price and stock.</li>\n          </ul>\n          <p>It doesn't replace a usable website. Many agents will keep using your pages for a long time, and the people they hand back to always will. Think of MCP and WebMCP as an extra, faster lane for the agents that support them.</p>\n          <h2>What MCP is</h2>\n          <p>The <a href=\"https://modelcontextprotocol.io/\">Model Context Protocol</a> (MCP) is an open protocol for connecting AI applications to outside tools and data. Anthropic introduced it in late 2024, and it has since been adopted by many AI assistants, developer tools and software companies.</p>\n          <p>An MCP server is a small service that describes a set of <strong>tools</strong> (actions an agent can call, each with a name, a plain-language description and the inputs it expects) and can also offer <strong>resources</strong> (data it can read, such as a catalog or a policy). An AI application connects to the server, reads the list, and calls the tools it needs to complete a person's request.</p>\n          <p>For a website, a remote MCP server runs alongside your site, usually on your own domain, and talks to the same systems your site does. Typical tools for an online store:</p>\n          <table>\n            <thead><tr><th>Tool</th><th>What it does</th><th>Risk</th></tr></thead>\n            <tbody>\n              <tr><td><code>search_products</code></td><td>Finds products by keyword, category, size or price</td><td>Read-only</td></tr>\n              <tr><td><code>get_product</code></td><td>Returns price, variants, stock and delivery estimate for one product</td><td>Read-only</td></tr>\n              <tr><td><code>get_policy</code></td><td>Returns shipping, returns or warranty terms</td><td>Read-only</td></tr>\n              <tr><td><code>get_order_status</code></td><td>Looks up an order for a signed-in customer</td><td>Needs sign-in</td></tr>\n              <tr><td><code>add_to_cart</code></td><td>Builds a cart and returns a link for the person to check out</td><td>Changes state</td></tr>\n            </tbody>\n          </table>\n          <p>Software companies might offer tools to look up plans and pricing or search documentation. Service businesses might offer tools to check availability and request a booking.</p>\n          <p>Commerce platforms are starting to offer MCP servers for the stores they host, so check with yours before building one. If you have a developer team and an existing API, a read-only server is often a small project, because it wraps calls you already make.</p>\n          <h2>What WebMCP is</h2>\n          <p>WebMCP is an early proposal for bringing the same idea into the web page itself. Instead of running a separate server, your page registers tools with the browser through a new JavaScript API, <code>navigator.modelContext</code>. An agent working in that browser can then call the page's tools directly instead of clicking through it.</p>\n          <p>It's being developed in the open at the W3C's Web Machine Learning Community Group, with engineers from Google and Microsoft among those working on it. Community group work is incubation, not a finished standard, and browsers don't ship WebMCP by default yet. Expect the details to change.</p>\n          <p>Two features make it interesting for websites:</p>\n          <ul>\n            <li><strong>It reuses your page code.</strong> The tool can call the same functions your search box or add-to-cart button already calls.</li>\n            <li><strong>It works with the person's session.</strong> Because it runs in the page, it can use the cart and sign-in the person already has, without separate API credentials.</li>\n          </ul>\n          <p>A tool registered in a page looks roughly like this, guarded so it does nothing in browsers without the API:</p>\n<pre><code>if (\"modelContext\" in navigator) {\n  navigator.modelContext.registerTool({\n    name: \"search_products\",\n    description: \"Search Northwind Outdoor products by keyword, size and maximum price.\",\n    inputSchema: {\n      type: \"object\",\n      properties: {\n        query: { type: \"string\" },\n        size: { type: \"string\" },\n        maxPrice: { type: \"number\" }\n      },\n      required: [\"query\"]\n    },\n    async execute({ query, size, maxPrice }) {\n      const results = await searchCatalog({ query, size, maxPrice }); // your existing search\n      return { content: [{ type: \"text\", text: JSON.stringify(results) }] };\n    }\n  });\n}</code></pre>\n          <p>The proposal also describes a declarative form: adding a <code>toolname</code> attribute (with a description) to an existing <code>&lt;form&gt;</code>, so a search or contact form becomes a tool without new JavaScript. Attribute names may change as the proposal develops.</p>\n          <h2>MCP or WebMCP?</h2>\n          <table>\n            <thead><tr><th></th><th>MCP server</th><th>WebMCP</th></tr></thead>\n            <tbody>\n              <tr><td>Where it runs</td><td>A service on your domain</td><td>Inside your web pages</td></tr>\n              <tr><td>Who can use it</td><td>AI applications that connect to MCP servers</td><td>Agents working in a browser that supports it</td></tr>\n              <tr><td>Maturity</td><td>Published specification, widely used</td><td>Early proposal, not shipped by default</td></tr>\n              <tr><td>Sign-in</td><td>Needs its own authorization for personal data</td><td>Uses the person's existing session</td></tr>\n              <tr><td>Best first step</td><td>Read-only catalog, pricing and policy tools</td><td>Mark up your search form, then experiment</td></tr>\n            </tbody>\n          </table>\n          <p>For most businesses, an MCP server is the practical choice today, and WebMCP is worth a small experiment so you're ready if browsers adopt it.</p>\n          <h2>What AgentScore looks for</h2>\n          <p>The <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> check \"Agents can use your site through MCP or WebMCP\" is part of the Access category. It passes if it finds any of these:</p>\n          <ul>\n            <li><strong>An MCP server description</strong> at <code>/.well-known/mcp.json</code> or <code>/.well-known/mcp</code>. There's no settled standard for advertising an MCP server yet, and proposals for a well-known server description are still being discussed, so we look in the most likely places.</li>\n            <li><strong>A link in your home page</strong>: a <code>&lt;link&gt;</code> or <code>&lt;a&gt;</code> whose <code>rel</code> or <code>type</code> mentions MCP.</li>\n            <li><strong>An MCP server listed in your llms.txt</strong>, with its URL. See <a href=\"https://ghostagentlab.com/articles/llms-txt/\">how to write an llms.txt file</a>.</li>\n            <li><strong>WebMCP tools registered in your home page.</strong> Because browsers don't ship WebMCP yet, the scanner provides a stand-in <code>navigator.modelContext</code> before your scripts run, so pages that check for the API register their tools as they would in a supporting browser. We record the tool names.</li>\n            <li><strong>Forms marked up as WebMCP tools</strong> with a <code>toolname</code> attribute.</li>\n          </ul>\n          <p>If none is found, the check is a warning rather than a failure. Agents can still use your pages; they just have to do it the slow way. It carries less weight than the checks that decide whether agents can get in at all, such as bot protection and robots.txt.</p>\n          <div>\n            <p><strong>Advertising a server isn't the same as it working.</strong> AgentScore confirms you've told agents where to find your tools. It doesn't call them. Test your tools with an MCP client before you publish the link, and keep them working as your catalog changes.</p>\n          </div>\n          <h2>How to start, safely</h2>\n          <ol>\n            <li><strong>Ask your platform first.</strong> If your commerce platform or CMS offers an MCP server or WebMCP support, turning it on is far cheaper than building one.</li>\n            <li><strong>Start read-only.</strong> Search, product details, pricing, stock and policies. These answer most agent questions and can't change anything.</li>\n            <li><strong>Write tool descriptions for a reader who has never seen your site.</strong> The description is how an agent decides which tool to call. \"Search products by keyword, size and maximum price; returns name, price, stock and URL\" beats \"Search\".</li>\n            <li><strong>Return the same facts your pages show.</strong> Prices, stock and policies must match the website exactly, or agents will quote figures your checkout won't honor.</li>\n            <li><strong>Keep people in charge of money.</strong> For anything that changes state, such as building a cart, return a link the person opens to review and pay. Leave fully agent-led checkout to the <a href=\"https://ghostagentlab.com/articles/agentic-commerce-protocols/\">agentic commerce protocols</a> designed for it.</li>\n            <li><strong>Protect it like an API.</strong> Rate limits, logging and authorization for anything personal. See <a href=\"https://ghostagentlab.com/articles/rate-limits-ai-agents/\">rate limits for AI agents</a>.</li>\n            <li><strong>Advertise it.</strong> List the server in llms.txt with a link to its documentation, and add a well-known description or a <code>&lt;link&gt;</code> in your page head.</li>\n          </ol>\n          <h2>What it means for the business</h2>\n          <p>An MCP server or WebMCP tools won't bring traffic on their own. They make each agent visit more likely to end well: the right product found, the right price quoted, the cart built. Because support among AI applications is still growing, the best order of work is the usual one. Make sure agents can get in and read your pages first, then add a direct lane for the agents that can use it.</p>\n          <p>Not sure where you stand on the basics? Start with <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">what agent readiness means</a>, then run AgentScore to see every Access check, including this one.</p>",
      "date_published": "2026-10-09T12:14:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/agentic-commerce-protocols/",
      "url": "https://ghostagentlab.com/articles/agentic-commerce-protocols/",
      "title": "Agentic commerce protocols explained",
      "summary": "What ACP, UCP and AP2 are for, what they mean for your store, what to ask your platform, and what AgentScore can and can't detect.",
      "content_html": "<p>Agentic commerce protocols let an AI assistant buy from your store on a person's behalf without filling in your checkout pages. The person says \"yes, buy it\" in the assistant, and the order arrives in your systems like any other. Several protocols have been announced in a short time, by different companies, and the names are easy to mix up. This guide explains what each one is for, what it means for your store, and what to ask your platform.</p>\n          <h2>Why protocols, when agents can use checkout pages?</h2>\n          <p>A browser agent can shop through your website, and for now that's how most agents buy. We cover that journey in <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">how to make checkout work for AI shopping agents</a>. But checkout pages were built for people. An agent filling them in has to guess at fields, cope with pop-ups and CAPTCHAs, and usually hand back to the person to pay.</p>\n          <p>A protocol replaces the page-filling with a structured conversation between the assistant and your store:</p>\n          <ul>\n            <li><strong>The assistant asks</strong> for a product, variant, quantity and delivery address.</li>\n            <li><strong>Your store answers</strong> with the real total, shipping options, tax and any errors, in data rather than on a page.</li>\n            <li><strong>The person approves</strong> the purchase inside the assistant.</li>\n            <li><strong>Payment is passed</strong> in a form that proves the person authorized it, without the assistant seeing raw card numbers.</li>\n            <li><strong>Your store creates the order</strong> and stays responsible for fulfillment, returns and customer service.</li>\n          </ul>\n          <h2>The protocols you'll hear about</h2>\n          <p>This is a young and fast-moving area. The descriptions below are deliberately brief, and the protocols' own documentation is the place for detail. Programs, eligibility and regions change often.</p>\n          <table>\n            <thead><tr><th>Protocol</th><th>Who's behind it</th><th>What it covers</th></tr></thead>\n            <tbody>\n              <tr><td>Agentic Commerce Protocol (ACP)</td><td>OpenAI and Stripe</td><td>Checkout between an AI assistant and a merchant. It's the protocol behind buying inside ChatGPT, and it has been published as an open specification.</td></tr>\n              <tr><td>Universal Commerce Protocol (UCP)</td><td>Google, with industry partners</td><td>A broader standard for agents to discover what a store supports and complete checkout. A store publishes a profile at <code>/.well-known/ucp</code>.</td></tr>\n              <tr><td>Agent Payments Protocol (AP2)</td><td>Google, with payments industry partners</td><td>The payment step specifically: a way to prove that a person authorized an agent to make a purchase, designed to work with different payment methods and alongside other protocols.</td></tr>\n            </tbody>\n          </table>\n          <p>Card networks and payment providers have also announced their own programs for payments made by agents. Most of these sit underneath the protocols above, at the payment layer, rather than competing with them.</p>\n          <div>\n            <p><strong>Checkout and payment are different jobs.</strong> ACP and UCP are about the shopping conversation and creating the order. AP2 is about trusting the payment. A store may end up supporting more than one, often without knowing, because the platform and payment provider handle it.</p>\n          </div>\n          <h2>What it means for your store</h2>\n          <ul>\n            <li><strong>Assistants become a sales channel.</strong> Supporting a protocol can let a person buy from you without leaving the assistant they asked. That puts you in front of shoppers at the moment they decide.</li>\n            <li><strong>You stay the merchant.</strong> In the protocols announced so far, the store keeps the order, the customer relationship and responsibility for fulfillment. Check the terms of any program you join.</li>\n            <li><strong>Your data has to be right.</strong> An assistant will quote the price, stock and delivery date your feed or API returns. Mistakes become disappointed customers at scale. See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data AI shopping agents can read</a>.</li>\n            <li><strong>Fraud and risk rules need a second look.</strong> Orders arrive without the usual browser signals. Ask your payment provider how it scores them.</li>\n            <li><strong>Returns and support still come to you.</strong> Make sure your team can see which orders came through an assistant.</li>\n          </ul>\n          <h2>How stores get support</h2>\n          <p>Very few stores will build a protocol integration themselves. Support usually arrives in one of three ways:</p>\n          <ol>\n            <li><strong>Through your commerce platform.</strong> Hosted platforms are adding protocol support for the stores they run, sometimes as a setting you switch on.</li>\n            <li><strong>Through your payment provider.</strong> Some providers handle the agent-facing checkout and payment for you.</li>\n            <li><strong>Through an AI platform's merchant program.</strong> Some assistants ask merchants to apply, and review product data and policies before listing them.</li>\n          </ol>\n          <p>If you run a custom stack, the specifications are public, but budget for real engineering: order creation, tax, shipping, inventory and payment all have to behave exactly as they do on your site.</p>\n          <h2>Questions to ask your platform and payment provider</h2>\n          <ul>\n            <li>Which agentic commerce protocols do you support today, and which are planned?</li>\n            <li>Is it on by default, or do we need to enable it or apply?</li>\n            <li>Which product data does it use, and how do we keep it in sync with the website?</li>\n            <li>How are agent orders marked in our order system and analytics?</li>\n            <li>How are fraud screening, chargebacks and refunds handled for agent orders?</li>\n            <li>Can we limit it to certain products, countries or order values while we learn?</li>\n          </ul>\n          <h2>What AgentScore checks</h2>\n          <p>For stores, <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> includes the Access check \"Agents can check out through an agentic commerce protocol\". It's skipped for sites that aren't stores.</p>\n          <p>It can only see what a store publishes openly. Today that means a Universal Commerce Protocol profile: the scanner requests <code>/.well-known/ucp</code>, and if it gets back a JSON document, the check passes.</p>\n          <p>The Agentic Commerce Protocol is different. Integrations are set up privately between each store and each AI platform, so there's nothing on your site for a scanner to find. If you support ACP but not UCP, this check will still show a warning. That's a limit of what can be seen from outside, not a problem with your store. The advice in the report names both protocols.</p>\n          <p>A warning here costs little. The check has a small weight in the Access category, because most agents still shop through the website, and that's where the checks that decide whether agents can get in and get through checkout matter more. Those include bot protection, CAPTCHAs at the cart and checkout, and <a href=\"https://ghostagentlab.com/articles/guest-checkout-ai-agents/\">guest checkout</a>.</p>\n          <h2>Where to start</h2>\n          <ol>\n            <li><strong>Fix the website journey first.</strong> Make sure agents can get in, read your product pages and reach checkout. Every protocol relies on the same product data and policies.</li>\n            <li><strong>Get your product data in order.</strong> Accurate prices, variants, stock and delivery times in structured data and in any product feeds you publish.</li>\n            <li><strong>Ask the questions above</strong> of your platform and payment provider, and turn on what's available.</li>\n            <li><strong>Watch how agents arrive.</strong> Track agent visits and the orders that come from assistants, so you can tell whether a protocol is earning its keep. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure AI agent traffic</a> and <a href=\"https://ghostagentlab.com/articles/ai-referral-traffic/\">tracking visits and sales from AI assistants</a>.</li>\n          </ol>\n          <p>Agentic commerce protocols will keep changing for a while. The stores best placed to benefit are the ones whose websites, data and policies already work for agents, because that's what every protocol is built on.</p>",
      "date_published": "2026-10-09T12:13:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/guest-checkout-ai-agents/",
      "url": "https://ghostagentlab.com/articles/guest-checkout-ai-agents/",
      "title": "Guest checkout: why AI shopping agents need it",
      "summary": "Why forced accounts stop AI agents at the last step, how AgentScore checks for guest checkout, and how to offer it without losing accounts.",
      "content_html": "<p>An AI agent can find your product, choose the size and fill the cart, then stop dead at a page that says \"Sign in to continue\". It can't create an account for the person it's shopping for, and it usually can't sign in as them either. Guest checkout removes that dead end. This guide explains why it matters, how AgentScore checks for it, and how to offer it without giving up on customer accounts.</p>\n          <h2>Why accounts stop agents</h2>\n          <p>A forced account is a known source of abandoned checkouts with people. With AI agents it's usually worse, for practical reasons:</p>\n          <ul>\n            <li><strong>Agents can't sign up on someone's behalf.</strong> Creating an account means choosing a password, agreeing to terms and often confirming an email address. Well-behaved agents won't invent a password or accept terms for a person, and they can't open the person's inbox to click a confirmation link.</li>\n            <li><strong>Signing in is a handoff.</strong> Browser agents generally stop at sign-in and pass control back to the person. Some people will finish; many won't, especially if they asked the assistant precisely so they wouldn't have to.</li>\n            <li><strong>Sign-in pages attract bot defenses.</strong> Login forms are a favorite target for credential-stuffing attacks, so they often carry the strictest CAPTCHAs and bot rules. An agent sent there meets the hardest wall on your site.</li>\n            <li><strong>The failure happens at the last step.</strong> The agent has already done the expensive work: searching, comparing, choosing. Losing the order at sign-in wastes all of it, and the assistant may recommend a store where it could finish instead.</li>\n          </ul>\n          <p>Guest checkout fixes all four. The agent, or the person it hands back to, enters an email address and a delivery address and carries on to payment.</p>\n          <h2>What good guest checkout looks like</h2>\n          <ul>\n            <li><strong>Guest is a clear choice, in words.</strong> A button or link that says \"Check out as guest\" or \"Continue as guest\", not an unlabelled icon or a link hidden below a sign-in form.</li>\n            <li><strong>Guest is the default path</strong>, or at least equal to sign-in. Many stores go straight to the contact and delivery form and offer sign-in as a link at the top.</li>\n            <li><strong>It asks only for what the order needs:</strong> email, name, delivery address, phone if your carrier requires it. Every field is labelled and uses the right <code>autocomplete</code> value. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>.</li>\n            <li><strong>No password field</strong> on the guest path. If you want to offer an account, offer it after the order.</li>\n            <li><strong>No CAPTCHA on every checkout.</strong> Use risk-based checks that challenge only suspicious orders. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a>.</li>\n            <li><strong>The cart doesn't say otherwise.</strong> A cart page that reads \"Sign in to check out\" undoes a perfectly good guest checkout behind it.</li>\n          </ul>\n          <h2>Keeping the benefits of accounts</h2>\n          <p>Accounts matter to most stores: repeat purchases, order history, loyalty and marketing consent. Guest checkout doesn't mean giving those up. It means moving the ask to a moment when it doesn't cost you the sale.</p>\n          <table>\n            <thead><tr><th>Instead of</th><th>Try</th></tr></thead>\n            <tbody>\n              <tr><td>Requiring an account before checkout</td><td>Offering \"Save your details for next time\" on the order confirmation page</td></tr>\n              <tr><td>Asking for a password at checkout</td><td>Emailing a link to set up an account from the order</td></tr>\n              <tr><td>Hiding order tracking behind sign-in</td><td>Order tracking by order number and email address</td></tr>\n              <tr><td>Requiring sign-in for loyalty points</td><td>Matching guest orders to existing accounts by email, then awarding points</td></tr>\n              <tr><td>An account wall for members-only prices</td><td>Showing the public price to guests and the member price once signed in</td></tr>\n            </tbody>\n          </table>\n          <p>Some stores genuinely need accounts: trade-only wholesale, age-restricted products, prescription items, or subscription services where the account is the product. If that's you, say so clearly on product pages and in your <a href=\"https://ghostagentlab.com/articles/llms-txt/\">llms.txt</a>, so agents can tell the person up front rather than discovering it at checkout.</p>\n          <h2>How AgentScore checks for guest checkout</h2>\n          <p>The Access check \"Shoppers can check out without an account\" applies to stores. The scanner is read-only: it loads your cart and checkout pages, but never submits a form or adds anything to a cart. It looks at what a shopper would see:</p>\n          <ul>\n            <li><strong>Pass</strong> if the cart or checkout offers a guest option in words, such as \"Check out as guest\", \"Continue as guest\" or \"No account needed\", or if the checkout asks for contact and delivery details with no password field.</li>\n            <li><strong>Fail</strong> if the checkout redirects to a sign-in page, asks for a password with no guest option, or the cart says shoppers must sign in to check out.</li>\n            <li><strong>Not tested</strong> if the checkout only opens with items in the cart. Because the scanner never adds items, many checkouts send it back to an empty cart. In that case the check is skipped and doesn't count against your score, rather than guessing.</li>\n          </ul>\n          <p>It carries a moderate weight in the Access category, alongside the related check that AI agents aren't blocked or shown a CAPTCHA at the cart and checkout. If your checkout couldn't be tested, a Ghost Agent can walk the full journey with items in the cart. See <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">how to test your key journeys with AI agents</a>.</p>\n          <h2>How to check your own store in five minutes</h2>\n          <ol>\n            <li>Open your store in a private browsing window, so you're signed out.</li>\n            <li>Add any product to the cart and open the cart page. Read it as an agent would: does any text say you must sign in?</li>\n            <li>Press the checkout button. Note where you land: a sign-in page, a choice between sign-in and guest, or straight into a contact and delivery form.</li>\n            <li>If there's a guest option, make sure it's a real button or link with words on it, and that it works with the keyboard alone.</li>\n            <li>Go through to the payment step without creating an account. Note every field, pop-up and challenge on the way.</li>\n          </ol>\n          <div>\n            <p><strong>Platform settings often decide this.</strong> Most commerce platforms have a single setting for whether accounts are required, optional or disabled at checkout. If yours is set to required, changing it may be the quickest win on your whole AgentScore report. Check with your platform admin before redesigning anything.</p>\n          </div>\n          <h2>Guest checkout and agentic commerce</h2>\n          <p>Assistants that buy through an <a href=\"https://ghostagentlab.com/articles/agentic-commerce-protocols/\">agentic commerce protocol</a> don't use your checkout pages, so they bypass this problem. But most agents still shop through the website, and the people they hand back to always do. Guest checkout is the simplest change that makes both work, and it helps human shoppers in a hurry just as much.</p>\n          <p>For the rest of the journey, from search to the payment handoff, see <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">how to make checkout work for AI shopping agents</a>. Then run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> to see where your store stands.</p>",
      "date_published": "2026-10-09T12:12:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/rate-limits-ai-agents/",
      "url": "https://ghostagentlab.com/articles/rate-limits-ai-agents/",
      "title": "Rate limits for AI agents: fair limits that don't turn customers away",
      "summary": "How to set rate limits that stop scrapers but not AI assistants: 429 with Retry-After, per-agent limits for verified agents, and what to watch.",
      "content_html": "<p>Rate limits protect your site from scrapers and floods of automated traffic. Set badly, they also turn away the AI assistants that are looking up your products for a customer. This guide explains how to set limits that stop abuse without blocking the agents you want, and how to say \"slow down\" in a way agents understand.</p>\n          <h2>Why limits catch good agents</h2>\n          <p>Most rate limits count requests per IP address. That made sense when one address meant one visitor. It works less well for AI agents:</p>\n          <ul>\n            <li><strong>Agents share addresses.</strong> Assistant fetchers and browser agents run in cloud data centers, so many different people's requests can arrive from a small pool of addresses. A per-IP limit treats them as one heavy user.</li>\n            <li><strong>Agent visits come in bursts.</strong> A person asks an assistant to compare three jackets, and it opens a category page, three product pages and a returns policy within seconds. That's a normal visit, but it can look like a spike.</li>\n            <li><strong>Crawlers and assistants get lumped together.</strong> A training crawler reading your whole catalog and an assistant answering one shopper's question have very different value to you, but a single \"bots\" limit treats them the same.</li>\n            <li><strong>Blocks look like limits.</strong> Many setups answer a rate-limited request with a 403 Forbidden or a challenge page. The agent can't tell it should wait, so it reports your site as unavailable.</li>\n          </ul>\n          <h2>Say \"slow down\" properly: 429 and Retry-After</h2>\n          <p>HTTP has a status code made for this. <strong>429 Too Many Requests</strong>, defined in <a href=\"https://www.rfc-editor.org/rfc/rfc6585\">RFC 6585</a>, tells the client it has sent too many requests. Pair it with a <strong>Retry-After</strong> header, defined in <a href=\"https://www.rfc-editor.org/rfc/rfc9110\">RFC 9110</a>, saying how many seconds to wait:</p>\n<pre><code>HTTP/1.1 429 Too Many Requests\nRetry-After: 30\nContent-Type: text/plain\nToo many requests from this client. Please wait 30 seconds and try again.</code></pre>\n          <p>Well-behaved crawlers and agents treat this as an instruction, not a door slammed shut. Google's crawler documentation, for example, says Googlebot slows down when it gets 429, 500 or 503 responses. Compare the alternatives:</p>\n          <table>\n            <thead><tr><th>Response</th><th>What the agent concludes</th></tr></thead>\n            <tbody>\n              <tr><td>429 with Retry-After</td><td>Wait, then try again. The site is fine.</td></tr>\n              <tr><td>429 without Retry-After</td><td>Back off, but guess for how long.</td></tr>\n              <tr><td>503 Service Unavailable</td><td>The site is having problems. Fine for real outages, misleading for rate limits.</td></tr>\n              <tr><td>403 Forbidden</td><td>I'm not allowed here. Many agents give up and say so.</td></tr>\n              <tr><td>200 with a challenge or \"access denied\" page</td><td>Confusing. Some agents read the challenge text as your content.</td></tr>\n            </tbody>\n          </table>\n          <p>Keep the 429 lightweight: a short plain-text message, no heavy page, and never a CAPTCHA, which agents can't solve. There's also an IETF draft for <code>RateLimit</code> headers that tell clients their remaining allowance before they hit the limit. It's still a draft, so treat it as a bonus for API clients, not a replacement for Retry-After.</p>\n          <h2>Limit by who, not just by address</h2>\n          <p>The fairest limits are set per agent, based on verified identity. Most CDNs and bot management products can verify the major AI agents against their operators' published IP ranges, reverse DNS or signed requests (see <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>). Once you know who's asking, you can give each kind of visitor a limit that fits its job:</p>\n          <table>\n            <thead><tr><th>Visitor</th><th>What it's doing</th><th>Suggested approach</th></tr></thead>\n            <tbody>\n              <tr><td>Verified AI assistants (such as ChatGPT-User, Claude-User, Perplexity-User)</td><td>Fetching pages because a person asked</td><td>Generous limits that allow short bursts. Each request stands in for a customer.</td></tr>\n              <tr><td>Verified AI search crawlers (such as OAI-SearchBot)</td><td>Indexing pages so you appear in AI answers</td><td>A steady crawl rate. Use 429 with Retry-After to pace them, not blocks.</td></tr>\n              <tr><td>Verified AI training crawlers (such as GPTBot)</td><td>Collecting content for model training</td><td>Your business decision. Pace them, or opt out in robots.txt. See <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>.</td></tr>\n              <tr><td>Unverified traffic claiming an AI agent's name</td><td>Unknown, sometimes impersonation</td><td>Normal or stricter limits. Never give extra allowance to a name alone.</td></tr>\n              <tr><td>Everything else</td><td>People, browsers, other bots</td><td>Your existing limits, tuned so a busy shopper never hits them.</td></tr>\n            </tbody>\n          </table>\n          <p>The right numbers depend on your traffic and infrastructure, so we don't suggest specific figures. Start from what your logs show a normal agent visit looks like, then set limits comfortably above it.</p>\n          <h2>Limit actions more than pages</h2>\n          <p>Abuse mostly targets actions: logins, account sign-ups, discount codes, gift card balances, stock checks and search. Pages that only display information are cheap to serve, especially from a cache. Put your tightest limits on:</p>\n          <ul>\n            <li>Login, sign-up and password reset</li>\n            <li>Coupon, gift card and checkout submission endpoints</li>\n            <li>Search and filter endpoints that hit your database</li>\n            <li>Public APIs and any <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP server</a> you offer</li>\n          </ul>\n          <p>Keep product, category, pricing, policy and help pages loose, and cache them well. Those are the pages agents need to answer questions about you, and the ones that turn into sales.</p>\n          <h2>Other ways to reduce load</h2>\n          <ul>\n            <li><strong>Cache aggressively.</strong> A product page served from a CDN cache costs almost nothing, however often it's requested.</li>\n            <li><strong>Support conditional requests.</strong> <code>ETag</code> and <code>Last-Modified</code> headers let crawlers ask \"has this changed?\" and get a tiny 304 Not Modified response if not.</li>\n            <li><strong>Keep your sitemap accurate.</strong> Correct <code>lastmod</code> dates help crawlers revisit only what changed. See <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">XML sitemaps for AI agents</a>.</li>\n            <li><strong>Don't rely on Crawl-delay.</strong> This robots.txt line is ignored by Google and supported unevenly elsewhere. Rate limits with 429 work for every client.</li>\n          </ul>\n          <h2>How to tell if limits are hurting you</h2>\n          <ol>\n            <li><strong>Check your logs by agent.</strong> Look at the share of requests from AI assistants that get 429, 403 or challenge responses. A handful of 429s for a crawler is healthy; repeated 429s or 403s for assistant fetchers mean customers' questions are going unanswered. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure AI agent traffic</a>.</li>\n            <li><strong>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>.</strong> It requests your home page as a normal browser and as several AI agents, and requests your key pages, cart and checkout as AI assistants. Any error status an agent gets while the browser gets through, including 429, counts as blocked in the bot protection, key pages and checkout checks.</li>\n            <li><strong>Ask your CDN provider</strong> whether its rate limiting treats verified bots separately, and whether it returns 429 or 403 when a limit is hit.</li>\n          </ol>\n          <div>\n            <p><strong>A quick checklist.</strong> Rate-limited responses use 429 with Retry-After, never 403 or a CAPTCHA. Verified AI assistants have their own, more generous limits. Names alone earn nothing extra. The strictest limits sit on logins, coupons and search, not on product and policy pages. And someone checks the 429 rate by agent at least monthly.</p>\n          </div>\n          <p>Rate limits and bot protection work together. For the wider picture, read <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a> and, for the business decision about which agents to let in at all, <a href=\"https://ghostagentlab.com/blog/block-or-allow-ai-crawlers/\">should you block AI crawlers?</a></p>",
      "date_published": "2026-10-09T12:11:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/image-alt-text-ai-agents/",
      "url": "https://ghostagentlab.com/articles/image-alt-text-ai-agents/",
      "title": "Alt text for AI agents: describing product images",
      "summary": "Why AI agents rely on alt text, what AgentScore checks, and how to write useful descriptions for product photos across a whole catalog.",
      "content_html": "<p>Your product photos do a lot of selling. They show the color, the cut, the size in someone's hand. Most AI agents never see them. What they read instead is the alt text: the short description attached to each image in your HTML. If it's missing, the photo might as well not be there.</p>\n          <h2>Why images are a blind spot for agents</h2>\n          <p>An AI agent helping someone shop usually reads your page as text. It fetches the HTML, strips out the layout and works from what's left. Images arrive as file names and URLs, which tell it very little. <code>IMG_4471_final.jpg</code> doesn't say whether the jacket is navy or black.</p>\n          <p>Some agents do look at the page as a picture, taking screenshots and reading them with a vision model. They can make out a photo, but slowly, at a cost, and not always correctly. Text is faster, cheaper and unambiguous, so even those agents lean on it when it's there.</p>\n          <p>Alt text is the bridge. It's an attribute on the image tag, written for anyone who can't see the image: screen reader users, people on slow connections, search engines and now AI agents. Writing it well is the same work for all of them. We make that case in <a href=\"https://ghostagentlab.com/blog/accessibility-is-agent-readiness/\">Accessibility work is agent readiness work</a>.</p>\n          <h2>What AgentScore checks</h2>\n          <p>The \"Images have text descriptions\" check is part of the Readability category. It loads your home page in a real browser and looks at the images people can actually see, ignoring tiny ones under 32 pixels wide such as icons and tracking pixels. For each one it asks a simple question: does it have an <code>alt</code> attribute at all?</p>\n          <table>\n            <thead><tr><th>Result</th><th>What it means</th></tr></thead>\n            <tbody>\n              <tr><td>Pass</td><td>90% or more of visible images have alt text, or there are no content images on the page</td></tr>\n              <tr><td>Warning</td><td>Between 60% and 90% have alt text</td></tr>\n              <tr><td>Fail</td><td>Fewer than 60% have alt text</td></tr>\n            </tbody>\n          </table>\n          <p>An empty <code>alt=\"\"</code> counts as described. That's deliberate: it's the correct way to mark an image as decorative, and it tells agents and screen readers to skip it. The check can't judge whether your descriptions are any good, so that part is up to you.</p>\n          <p>It's a small check by weight, because a missing description rarely stops an agent on its own. But the same gaps tend to show up elsewhere. A linked product image with no alt text is also a link with no name, and that's caught by the \"Buttons and links have names\" check in Navigability, which carries more weight.</p>\n          <h2>How to write alt text for product images</h2>\n          <p>Describe what a shopper needs to know from the photo, in the words they'd use. Lead with the product, then add the details the image actually shows.</p>\n          <table>\n            <thead><tr><th>Weak</th><th>Better</th></tr></thead>\n            <tbody>\n              <tr><td><code>alt=\"product\"</code></td><td><code>alt=\"Northwind waxed cotton field jacket in navy, front view\"</code></td></tr>\n              <tr><td><code>alt=\"IMG_4471\"</code></td><td><code>alt=\"Field jacket in navy, close-up of brass zip and corduroy collar\"</code></td></tr>\n              <tr><td><code>alt=\"best jacket waterproof jacket mens jacket sale\"</code></td><td><code>alt=\"Model, 6 ft 1 in, wearing the field jacket in size L\"</code></td></tr>\n              <tr><td>No alt attribute</td><td><code>alt=\"\"</code> if the image is purely decorative</td></tr>\n            </tbody>\n          </table>\n          <ul>\n            <li><strong>Name the product and the variant.</strong> If the photo shows the navy version, say navy. Agents use this to match images to options.</li>\n            <li><strong>Say what's different about each shot.</strong> Front, back, detail, in use, on a model. Five images all called \"Field jacket\" waste the opportunity.</li>\n            <li><strong>Include facts the photo proves.</strong> Model height and size worn, what's in the box, the scale of an object next to something familiar.</li>\n            <li><strong>Keep it short.</strong> One short sentence is usually enough. Long specifications belong in the product description, not the image.</li>\n            <li><strong>Don't stuff keywords.</strong> Agents read alt text as a description. A list of search terms reads as noise and helps nobody.</li>\n            <li><strong>Don't start with \"image of\".</strong> Everyone reading it already knows it's an image.</li>\n          </ul>\n          <h2>Decorative images, logos and linked images</h2>\n          <p>Not every image needs words. Background textures, divider graphics and lifestyle photos that repeat the headline next to them can take <code>alt=\"\"</code>. That's different from leaving <code>alt</code> out: an empty value says \"nothing to see here\", while a missing one leaves agents and screen readers guessing, and some will read out the file name instead.</p>\n          <p>Images that are links need extra care, because the alt text becomes the link's name. A product tile where the whole image links to the product page should have alt text that works as a link, such as the product name. Your logo in the header links home, so <code>alt=\"Northwind home\"</code> or simply your brand name works better than <code>alt=\"logo\"</code>.</p>\n          <p>Text baked into images is a common trap on home pages. A banner that says \"40% off outerwear this weekend\" as pixels is invisible to most agents. Put the offer in real HTML text over the image, or at minimum repeat it in the alt text.</p>\n          <h2>Doing this at catalog scale</h2>\n          <p>A store with thousands of products can't write every description by hand. You don't need to. Most of it can be generated from data you already have.</p>\n          <ol>\n            <li><strong>Fix the template first.</strong> Make sure your theme outputs an <code>alt</code> attribute on every product image, and falls back to the product name and variant when no custom text is set. That alone gets most stores to a pass.</li>\n            <li><strong>Use the image metadata your platform stores.</strong> Shopify, WooCommerce, Magento and most other platforms have an alt text field per image. Fill it on new products as part of the listing checklist.</li>\n            <li><strong>Backfill your bestsellers.</strong> Write proper descriptions for the products that earn the most, where shoppers are most likely to ask an assistant about them.</li>\n            <li><strong>Review generated text before publishing.</strong> Tools that write alt text with AI are useful for a first draft, but they can describe the wrong color or invent details. Check a sample, especially for variants.</li>\n            <li><strong>Add it to the content workflow.</strong> Merchandisers and content teams usually own this, not developers. Make alt text a required field when new images are uploaded.</li>\n          </ol>\n          <div>\n            <p><strong>Check it yourself in a minute:</strong> open a product page, right-click a photo and choose Inspect. Look for <code>alt=\"…\"</code> on the <code>&lt;img&gt;</code> tag. If it's missing, or it just says \"image\", agents are getting nothing.</p>\n          </div>\n          <h2>Where alt text fits</h2>\n          <p>Alt text is the visual part of a bigger job: making sure everything a shopper learns from your page is also available as text. The other parts are prices and stock in a form agents can read (see <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">Prices AI agents can read</a> and <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">Product data that AI shopping agents can read</a>), and content that's in the HTML rather than added by JavaScript (see <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">our guide to JavaScript-only content</a>).</p>\n          <p>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> to see how your home page does, then spot-check a few product pages by hand. For the full picture of what agents need from a site, start with the <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness guide</a>.</p>",
      "date_published": "2026-10-09T12:10:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/",
      "url": "https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/",
      "title": "XML sitemaps for AI agents",
      "summary": "How AI agents and crawlers use your sitemap, what AgentScore checks, what to include and how to keep it accurate.",
      "content_html": "<p>A sitemap is a list of the pages on your site that you want found, in a format machines read without effort. Search engines have used them for about twenty years. AI agents and AI search crawlers use them too, to find pages your navigation hides and to discover new products quickly. It's one of the cheapest readiness wins there is.</p>\n          <h2>What a sitemap does for AI agents</h2>\n          <p>An agent arriving at your site has two ways to find things: follow links from the home page, or read a list someone has prepared. Following links is slow and fragile. Menus built with JavaScript, products buried four clicks deep and pages that only appear in search results can all be missed.</p>\n          <p>A sitemap solves that. It's an XML file, usually at <code>/sitemap.xml</code>, that lists your important URLs. Crawlers from AI search engines use it to decide what to fetch. Agents doing a task can use it to locate a product or policy page directly. And it gives every system the same, complete picture of what you publish.</p>\n          <p>It doesn't replace good links. A page that's only in your sitemap is still a page most visitors and many agents never reach. Treat the sitemap as the safety net under your navigation, not instead of it. Our guide to <a href=\"https://ghostagentlab.com/articles/link-key-pages-home-page/\">linking key pages from the home page</a> covers the other half.</p>\n          <h2>What AgentScore checks</h2>\n          <p>The \"Valid sitemap\" check, in the Readability category, looks for your sitemap the way crawlers do:</p>\n          <ol>\n            <li>It reads <code>Sitemap:</code> lines in your <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt</a>.</li>\n            <li>It then tries <code>/sitemap.xml</code> and <code>/sitemap_index.xml</code>.</li>\n            <li>For each, it checks the file loads with a 200 status and is a real sitemap: XML with a <code>&lt;urlset&gt;</code> or <code>&lt;sitemapindex&gt;</code> element.</li>\n          </ol>\n          <p>It passes when it finds one, and reports how many entries it holds and whether it's listed in robots.txt. It fails if none of those places has a valid sitemap. A home page or a \"not found\" page served at <code>/sitemap.xml</code> doesn't count, which catches a common problem with some site builders.</p>\n          <p>The sitemap also feeds other checks. AgentScore uses it to find a product or pricing page when the home page doesn't link to one. If the sitemap is an index, it reads product sitemaps first, since most platforms name them that way. That's why a missing sitemap can leave other checks marked \"not tested\": without a product page to look at, there's nothing to check prices or product data on.</p>\n          <h2>What a good sitemap looks like</h2>\n          <p>The format is defined at <a href=\"https://www.sitemaps.org/protocol.html\">sitemaps.org</a>. A minimal one for a small store:</p>\n          <pre><code>&lt;?xml version=\"1.0\" encoding=\"UTF-8\"?&gt;\n&lt;urlset xmlns=\"http://www.sitemaps.org/schemas/sitemap/0.9\"&gt;\n  &lt;url&gt;\n    &lt;loc&gt;https://northwind.example/&lt;/loc&gt;\n    &lt;lastmod&gt;2026-10-01&lt;/lastmod&gt;\n  &lt;/url&gt;\n  &lt;url&gt;\n    &lt;loc&gt;https://northwind.example/products/ethiopia-yirgacheffe&lt;/loc&gt;\n    &lt;lastmod&gt;2026-10-07&lt;/lastmod&gt;\n  &lt;/url&gt;\n  &lt;url&gt;\n    &lt;loc&gt;https://northwind.example/pages/shipping&lt;/loc&gt;\n    &lt;lastmod&gt;2026-08-14&lt;/lastmod&gt;\n  &lt;/url&gt;\n&lt;/urlset&gt;</code></pre>\n          <p>Larger sites split it up. Each sitemap file can hold up to 50,000 URLs, and a sitemap index lists the individual files:</p>\n          <pre><code>&lt;?xml version=\"1.0\" encoding=\"UTF-8\"?&gt;\n&lt;sitemapindex xmlns=\"http://www.sitemaps.org/schemas/sitemap/0.9\"&gt;\n  &lt;sitemap&gt;&lt;loc&gt;https://northwind.example/sitemap_products_1.xml&lt;/loc&gt;&lt;/sitemap&gt;\n  &lt;sitemap&gt;&lt;loc&gt;https://northwind.example/sitemap_collections_1.xml&lt;/loc&gt;&lt;/sitemap&gt;\n  &lt;sitemap&gt;&lt;loc&gt;https://northwind.example/sitemap_pages_1.xml&lt;/loc&gt;&lt;/sitemap&gt;\n&lt;/sitemapindex&gt;</code></pre>\n          <p>Then tell crawlers where it is, with one line in robots.txt:</p>\n          <pre><code>Sitemap: https://northwind.example/sitemap.xml</code></pre>\n          <h2>What to include, and what to leave out</h2>\n          <table>\n            <thead><tr><th>Include</th><th>Leave out</th></tr></thead>\n            <tbody>\n              <tr><td>Product and category pages you want people to find</td><td>Cart, checkout and account pages</td></tr>\n              <tr><td>Pricing, plans and sign-up pages</td><td>Internal search results and filtered URLs</td></tr>\n              <tr><td>Shipping, returns, FAQ and contact pages</td><td>Pages that redirect elsewhere</td></tr>\n              <tr><td>Guides, articles and other evergreen content</td><td>Pages marked noindex, or blocked in robots.txt</td></tr>\n              <tr><td>The canonical URL of each page</td><td>Duplicates with tracking parameters or session IDs</td></tr>\n            </tbody>\n          </table>\n          <p>The rule of thumb: list the pages you'd be happy for an AI assistant to send a customer to, each at its one true address. If you include a page in the sitemap but block it in robots.txt, you're sending crawlers mixed signals.</p>\n          <h2>Keep it accurate</h2>\n          <ul>\n            <li><strong>Generate it automatically.</strong> Shopify, WooCommerce, Magento, BigCommerce, WordPress SEO plugins and most modern frameworks produce a sitemap for you. A hand-made sitemap goes stale the week after it's written.</li>\n            <li><strong>Make <code>lastmod</code> honest.</strong> It should change when the content does, such as a price or stock change, not every time the file is rebuilt. <a href=\"https://developers.google.com/search/docs/crawling-indexing/sitemaps/overview\">Google's documentation</a> says it uses <code>lastmod</code> only when it's consistently accurate, and other crawlers are likely to treat it similarly.</li>\n            <li><strong>Don't bother with priority and changefreq.</strong> They're in the protocol, but Google says it ignores them. Your time is better spent on accurate dates.</li>\n            <li><strong>Remove dead pages.</strong> Discontinued products that return 404, or redirect to a category, shouldn't stay listed.</li>\n            <li><strong>Keep it reachable.</strong> Make sure bot protection doesn't challenge crawlers asking for the sitemap. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">our guide to bot protection and CAPTCHAs</a>.</li>\n          </ul>\n          <h2>Sitemap, llms.txt or both?</h2>\n          <p>Both, because they do different jobs. A sitemap is complete: every page you want found, with no explanation. An <a href=\"https://ghostagentlab.com/articles/llms-txt/\">llms.txt file</a> is selective: a short, described guide to the pages that matter most, written for AI. A crawler indexing your catalog wants the sitemap. An assistant trying to answer \"what's their returns window?\" benefits from llms.txt pointing straight at the returns page.</p>\n          <div>\n            <p><strong>Quick check:</strong> open <code>yourdomain.com/robots.txt</code> and look for a <code>Sitemap:</code> line. Then open the URL it gives. You should see XML listing your pages, not a web page. If either step fails, that's your first fix.</p>\n          </div>\n          <h2>Next steps</h2>\n          <ol>\n            <li>Confirm your platform generates a sitemap, and that it loads.</li>\n            <li>Add a <code>Sitemap:</code> line to robots.txt if it's missing.</li>\n            <li>Spot-check that your bestselling products and your policy pages are listed.</li>\n            <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>. The \"Valid sitemap\" check will show what it found and how many entries it counted.</li>\n          </ol>\n          <p>For how the sitemap fits with the other things agents need, see the <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness guide</a>.</p>",
      "date_published": "2026-10-09T12:09:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/titles-descriptions-headings-ai/",
      "url": "https://ghostagentlab.com/articles/titles-descriptions-headings-ai/",
      "title": "Titles, descriptions and headings AI agents understand",
      "summary": "How AI agents use page titles, meta descriptions, headings and the lang attribute, and how to write each so agents pick the right page.",
      "content_html": "<p>Before an AI agent reads a page in full, it reads the labels: the title, the description and the headings. They're how it decides whether a page answers the question it was asked, and how it finds the right part of a long page. They're also the parts of a page most often left on autopilot.</p>\n          <h2>Why the labels matter more to agents than to people</h2>\n          <p>A person landing on your page takes in the design at a glance: the hero image, the logo, the layout. An agent sees none of that. It sees text, in order, and leans heavily on the few pieces of text that say what the page is.</p>\n          <ul>\n            <li><strong>The <code>&lt;title&gt;</code></strong> is the page's name. It shows in browser tabs and search results, and it's often the first thing an agent records about a page it has visited.</li>\n            <li><strong>The meta description</strong> is a one or two sentence summary. Search engines and AI search tools may show it, quote it, or use it to decide whether a page is worth opening.</li>\n            <li><strong>Headings</strong> (<code>&lt;h1&gt;</code> to <code>&lt;h6&gt;</code>) are the outline. An agent looking for \"delivery times\" on a long FAQ uses them to jump to the right section.</li>\n            <li><strong>The <code>lang</code> attribute</strong> on the <code>&lt;html&gt;</code> tag says what language the page is in, so the agent reads and quotes it correctly.</li>\n          </ul>\n          <p>If you've done SEO work, most of this will be familiar. The difference is that an agent isn't ranking your page among ten others; it's trying to complete a task, and a vague label can mean it picks the wrong page or gives up. There's more on that shift in <a href=\"https://ghostagentlab.com/blog/seo-to-agent-readiness/\">SEO got you found. Agent readiness gets you chosen</a>.</p>\n          <h2>What AgentScore checks</h2>\n          <p>The \"Clear page title, description, and headings\" check, in the Readability category, reads your home page's HTML before any JavaScript runs and looks for four things:</p>\n          <table>\n            <thead><tr><th>Looks for</th><th>Counts as missing when</th></tr></thead>\n            <tbody>\n              <tr><td>A <code>&lt;title&gt;</code></td><td>There's no title tag, or it's empty</td></tr>\n              <tr><td>A meta description</td><td>There's no <code>&lt;meta name=\"description\"&gt;</code>, or it's empty</td></tr>\n              <tr><td>An <code>&lt;h1&gt;</code> heading</td><td>There's no h1 in the HTML (one added later by JavaScript doesn't count)</td></tr>\n              <tr><td>A <code>lang</code> attribute</td><td>The <code>&lt;html&gt;</code> tag has no lang</td></tr>\n            </tbody>\n          </table>\n          <p>All four present is a pass, and the result shows your title. One or two missing is a warning. It fails if there's no title at all, or if three or more are missing. The check confirms these things exist; it can't tell whether the wording is helpful. The rest of this guide is about that part.</p>\n          <h2>Writing titles agents can use</h2>\n          <p>A good title says what the page is and whose it is, specific enough to tell it apart from every other page on your site.</p>\n          <table>\n            <thead><tr><th>Page</th><th>Weak</th><th>Better</th></tr></thead>\n            <tbody>\n              <tr><td>Home page</td><td>Home</td><td>Northwind Coffee: specialty coffee beans and brewing gear</td></tr>\n              <tr><td>Product</td><td>Shop</td><td>Ethiopia Yirgacheffe whole bean coffee, 340 g | Northwind Coffee</td></tr>\n              <tr><td>Category</td><td>Collection</td><td>Pour-over coffee kits | Northwind Coffee</td></tr>\n              <tr><td>Policy</td><td>Info</td><td>Shipping rates and delivery times | Northwind Coffee</td></tr>\n            </tbody>\n          </table>\n          <ul>\n            <li><strong>Put the specific part first.</strong> The product or topic, then your brand.</li>\n            <li><strong>Make every title unique.</strong> If twenty pages share a title, an agent comparing them has nothing to go on.</li>\n            <li><strong>Include the variant or size when it matters.</strong> \"340 g\" or \"Pro plan\" is exactly what an agent needs to match a request.</li>\n            <li><strong>Skip the slogans.</strong> \"Where every cup tells a story\" says nothing about what's on the page.</li>\n          </ul>\n          <h2>Writing descriptions that summarize the page</h2>\n          <p>Write the meta description as the answer to \"what will I find here?\" One or two plain sentences: what the page offers, for whom, and anything that sets it apart, such as price range, delivery area or a free trial.</p>\n          <pre><code>&lt;meta name=\"description\" content=\"Light-roast single-origin coffee from\nYirgacheffe, Ethiopia, with notes of jasmine and lemon. Roasted to order,\n340 g bag, ships free in the US over $40.\"&gt;</code></pre>\n          <p>Avoid repeating the title, stacking keywords, or reusing one description across many pages. If a platform fills it in automatically from the first lines of the page, check that those lines are useful. Cookie notices and \"Free shipping this week\" banners make poor summaries.</p>\n          <h2>Using headings as an outline</h2>\n          <p>Headings should read like a table of contents. Someone reading only your headings should understand what the page covers and where each part is.</p>\n          <ul>\n            <li><strong>One h1 per page</strong> that names the page's subject, usually close to the title. The check only needs at least one, but more than one makes the page's subject less clear.</li>\n            <li><strong>h2s for main sections, h3s inside them.</strong> Don't skip levels or pick a heading level for its size; use CSS for size.</li>\n            <li><strong>Make headings descriptive.</strong> \"Delivery times by region\" beats \"Good to know\". On FAQ pages, the question itself makes a good heading.</li>\n            <li><strong>Use real heading tags.</strong> Bold text styled to look like a heading isn't one. Agents, and screen readers, can't see the styling.</li>\n            <li><strong>Keep headings in the HTML.</strong> If your page builder adds them with JavaScript, many agents won't see them. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">our guide to JavaScript-only content</a>.</li>\n          </ul>\n          <p>A product page outline might look like this:</p>\n          <pre><code>h1  Ethiopia Yirgacheffe whole bean coffee\n  h2  Tasting notes\n  h2  Brewing suggestions\n  h2  Shipping and returns\n  h2  Reviews</code></pre>\n          <h2>Don't forget the lang attribute</h2>\n          <p>It's one attribute, and it's often missing from custom themes: <code>&lt;html lang=\"en\"&gt;</code>. Use the right code for each language version, such as <code>lang=\"fr\"</code> or <code>lang=\"en-GB\"</code>. It helps agents interpret the page, translate it correctly and quote it in the right language. Screen readers use it to choose a voice.</p>\n          <h2>Fixing it across a whole site</h2>\n          <ol>\n            <li><strong>Start with templates.</strong> Titles, descriptions and headings on product, category and policy pages usually come from a theme template. Fix the pattern once and every page benefits.</li>\n            <li><strong>Feed templates real data.</strong> Product name, variant, category and brand make good titles. Short product descriptions make good meta descriptions.</li>\n            <li><strong>Hand-write the important ones.</strong> Your home page, top categories and policy pages deserve titles and descriptions written by a person.</li>\n            <li><strong>Crawl for duplicates.</strong> Most SEO crawlers report missing and duplicate titles, descriptions and h1s across the whole site.</li>\n            <li><strong>Re-run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a></strong> to confirm the home page passes.</li>\n          </ol>\n          <div>\n            <p><strong>Who owns this:</strong> usually the SEO or content team, with a developer for template changes and the lang attribute. It's rarely a big project, which is why it's worth doing first.</p>\n          </div>\n          <p>Clear labels are one part of a readable site. For the rest, see the <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness guide</a>.</p>",
      "date_published": "2026-10-09T12:08:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/machine-readable-prices/",
      "url": "https://ghostagentlab.com/articles/machine-readable-prices/",
      "title": "Prices AI agents can read",
      "summary": "How to show prices AI agents read correctly: in the server HTML, with a clear currency, labelled sale prices, variant prices and one source of truth.",
      "content_html": "<p>\"How much is it?\" is one of the first things a shopper asks an AI assistant about a product. If an agent can't find your price, or finds two prices and can't tell which is real, it either guesses or moves on to a store where the answer is clear. Here's how to make your prices unmistakable.</p>\n          <h2>How agents read a price</h2>\n          <p>Agents find prices in two places. The first is the text of the page: the \"$18.00\" a shopper sees next to the add-to-cart button. The second is structured data: a block of schema.org JSON-LD in the HTML that states the price, currency and stock as plain values. Good stores have both, and they agree.</p>\n          <p>Both have to be in the HTML your server sends. Many agents, and most AI search crawlers, read that HTML without running JavaScript. If your theme loads prices from an API after the page appears, those agents see a product with no price. To a shopper asking an assistant to compare three coffee grinders, that's a product that doesn't make the list.</p>\n          <h2>What AgentScore checks</h2>\n          <p>Two Readability checks cover prices.</p>\n          <table>\n            <thead><tr><th>Check</th><th>What it looks at</th></tr></thead>\n            <tbody>\n              <tr><td>Prices are in the page HTML</td><td>Your pricing page, or a product page if there's no pricing page. It reads the text of the HTML before JavaScript runs and looks for something that's clearly a price: an amount with a currency symbol ($, €, £, ¥, ₹) or a currency code such as USD, EUR, GBP, CAD or AUD. It passes if it finds one. It warns if prices only appear after JavaScript runs in a real browser, or if it can't find a price at all.</td></tr>\n              <tr><td>Product pages give price and stock in a form agents can read</td><td>A product page's HTML, for schema.org <code>Product</code> or <code>ProductGroup</code> data with an offer that gives a price, a currency and availability. It passes with all three, warns if any are missing, if there's only microdata, or if the data is only added by JavaScript, and fails if there's none. Stores only.</td></tr>\n            </tbody>\n          </table>\n          <p>AgentScore finds these pages from links on your home page, then from your <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">sitemap</a>, then a common address such as <code>/pricing</code>. If it can't find one, the checks are marked \"not tested\" rather than failed.</p>\n          <p>This guide focuses on the visible price and keeping everything consistent. For a full walk-through of product structured data, see <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">Product data that AI shopping agents can read</a>.</p>\n          <h2>Put the price in the HTML</h2>\n          <p>Open a product page, right-click and choose View Page Source (not Inspect, which shows the page after JavaScript). Search for the price. If it's there, agents can read it. If it isn't, the fix belongs to your developer or platform: render the price on the server, as part of the page, and let JavaScript update it afterwards if a shopper changes options.</p>\n          <pre><code>&lt;p class=\"price\"&gt;\n  &lt;span&gt;$18.00&lt;/span&gt; &lt;span&gt;USD&lt;/span&gt;\n&lt;/p&gt;</code></pre>\n          <p>For software and services, the same applies to your pricing page: each plan's price, billing period and what it includes should be readable text, not an image or a widget that loads later. If you don't publish prices, say so in plain words, for example \"Pricing is quoted per project. Contact sales for a quote.\" Then an agent can tell people how to get one, instead of reporting that it couldn't find any.</p>\n          <h2>Make the currency unambiguous</h2>\n          <p>\"$18\" means different things in the US, Canada and Australia. A person infers the currency from where they are. An agent may be working for someone elsewhere, fetching your page from a data center in another country.</p>\n          <ul>\n            <li><strong>Use a currency code or prefix when it could be unclear.</strong> \"US$18.00\", \"$18.00 USD\" or \"CA$24.00\" leaves no doubt.</li>\n            <li><strong>Always set <code>priceCurrency</code> in structured data</strong> using the three-letter ISO 4217 code, such as <code>USD</code> or <code>EUR</code>.</li>\n            <li><strong>Be careful with automatic currency switching.</strong> If you change the currency based on the visitor's location, an agent may see a different currency from its user. Give each currency a stable URL or a visible switcher, and keep the currency in the page text and the data consistent.</li>\n            <li><strong>Say whether tax is included.</strong> \"Includes VAT\" or \"plus sales tax\" changes the answer to \"how much is it?\".</li>\n          </ul>\n          <h2>Sale prices and \"from\" prices</h2>\n          <p>Discounts are where agents most often misread prices. A page showing \"$24.00 $18.00\" with the first one struck through looks obvious to a person, because they can see the line through it. An agent reading text may see two prices side by side.</p>\n          <ul>\n            <li><strong>Label both prices in words.</strong> \"Was $24.00, now $18.00\" reads correctly in any form.</li>\n            <li><strong>Mark the old price up as removed.</strong> Use <code>&lt;del&gt;</code> or <code>&lt;s&gt;</code> for the original price, not just a CSS line-through.</li>\n            <li><strong>Put the price the customer will pay in structured data.</strong> The offer's <code>price</code> should be the current sale price. If the sale ends on a known date, <code>priceValidUntil</code> says so.</li>\n            <li><strong>Make \"from\" prices explicit.</strong> \"From $18.00\" on a product with several sizes should say what the $18.00 buys, such as \"From $18.00 (340 g)\".</li>\n          </ul>\n          <pre><code>&lt;p class=\"price\"&gt;\n  Was &lt;del&gt;$24.00&lt;/del&gt;, now &lt;strong&gt;$18.00 USD&lt;/strong&gt;\n  &lt;span&gt;Sale ends October 31&lt;/span&gt;\n&lt;/p&gt;</code></pre>\n          <h2>Products with variants</h2>\n          <p>When sizes, colors or pack sizes have different prices, agents need to know which price goes with which option. A single price that changes with JavaScript when a shopper picks a size is invisible to an agent that doesn't click.</p>\n          <ul>\n            <li><strong>List variant prices in the HTML</strong>, for example in the option labels: \"340 g, $18.00\" and \"1 kg, $44.00\".</li>\n            <li><strong>Describe each variant in structured data.</strong> Schema.org's <code>ProductGroup</code> with <code>hasVariant</code> gives each option its own offer with its own price and availability. AgentScore reads offers inside variants too.</li>\n            <li><strong>Give variants their own URLs</strong> where your platform supports it, so an agent can link a shopper to exactly the option they asked about.</li>\n          </ul>\n          <h2>Keep every price consistent</h2>\n          <p>An agent may see your price in several places: the page text, the JSON-LD, your product feed in Google Merchant Center or other shopping channels, and the cart. If they disagree, the agent can't know which is right. It may quote the wrong one, or treat your store as unreliable. Search engines may ignore structured data that doesn't match the page.</p>\n          <div>\n            <p><strong>One source of truth:</strong> generate the visible price, the structured data and your product feeds from the same data, so a price change updates all of them together. Hand-entered prices in a separate feed or a hard-coded JSON-LD block are the usual cause of mismatches.</p>\n          </div>\n          <p>Unit prices help too. For groceries, coffee, cosmetics and anything sold by weight or volume, a unit price such as \"$5.29 per 100 g\" lets an agent compare fairly across pack sizes.</p>\n          <h2>Your checklist</h2>\n          <ol>\n            <li>View the source of a product page and a pricing page. Is the price there?</li>\n            <li>Is the currency clear without knowing where the visitor is?</li>\n            <li>Do sale prices say which price is current, in words?</li>\n            <li>Does each variant's price appear in the HTML or the structured data?</li>\n            <li>Do the page, the JSON-LD, the feed and the cart all show the same price?</li>\n            <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> and check both price checks pass.</li>\n          </ol>\n          <p>Prices are one of the facts agents most need to get right before they recommend you or buy from you. For the step after that, getting the item into a basket and through checkout, see <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">our guide to agent-ready checkout</a>.</p>",
      "date_published": "2026-10-09T12:07:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/policy-pages-ai-assistants/",
      "url": "https://ghostagentlab.com/articles/policy-pages-ai-assistants/",
      "title": "Shipping, returns and FAQ pages AI assistants can quote",
      "summary": "How to write shipping, returns and FAQ pages that AI assistants can find, read and quote accurately to shoppers before they buy.",
      "content_html": "<p>\"Do they ship to Canada?\" \"Can I return it if it doesn't fit?\" \"How long does delivery take?\" These are the questions people ask AI assistants just before they buy. The answer comes from your shipping, returns and FAQ pages, quoted word for word or summarized. If those pages are vague, hidden or contradictory, the assistant gets it wrong, or tells the shopper it isn't sure.</p>\n          <h2>Why policy pages matter more now</h2>\n          <p>Policy pages used to be read by a small share of careful shoppers, and by your support team. Now an assistant may read them on behalf of every shopper who asks a question about delivery or returns. Whatever it says becomes your answer, with your name on it.</p>\n          <p>That has two consequences. A clear policy page means accurate answers and confident shoppers. A muddled one means an assistant may misquote your returns window, invent a shipping rate, or decline to answer and suggest a store whose terms it could read. Neither shows up in your analytics as an error. It shows up as sales that didn't happen, or customers who expected something you never promised.</p>\n          <p>AgentScore doesn't have a dedicated check for policy pages yet, so this guide is about what you can do and test yourself.</p>\n          <h2>Write so each sentence can be quoted on its own</h2>\n          <p>Assistants often lift a sentence or two out of a page. Write so that any sentence still makes sense when it's quoted without its neighbors.</p>\n          <table>\n            <thead><tr><th>Hard to quote</th><th>Easy to quote</th></tr></thead>\n            <tbody>\n              <tr><td>We're happy to help with returns within the usual period, subject to the conditions above.</td><td>You can return unopened items within 30 days of delivery for a full refund.</td></tr>\n              <tr><td>Shipping is fast and affordable.</td><td>US orders ship in 1 to 2 business days. Standard delivery costs $5.95 and takes 3 to 5 business days. Orders over $40 ship free.</td></tr>\n              <tr><td>International shipping may be available.</td><td>We ship to the US and Canada. We don't ship to other countries yet.</td></tr>\n              <tr><td>See exceptions.</td><td>Opened coffee and gift cards can't be returned.</td></tr>\n            </tbody>\n          </table>\n          <ul>\n            <li><strong>Use numbers, not adjectives.</strong> \"30 days\", \"$5.95\", \"3 to 5 business days\". \"Fast\" and \"generous\" can't be quoted usefully.</li>\n            <li><strong>Name the thing in every sentence.</strong> \"Returns are free\" works alone. \"They're free\" doesn't.</li>\n            <li><strong>State the conditions next to the rule.</strong> If the 30 days starts from delivery, not from the order date, say so in the same sentence.</li>\n            <li><strong>Say what you don't do.</strong> Countries you don't ship to, items that can't be returned, services you don't offer. Assistants are often asked exactly these questions, and silence leaves them guessing.</li>\n          </ul>\n          <h2>Structure the page around real questions</h2>\n          <p>Organize each page by the questions customers actually ask. Your support inbox and chat logs are the best source. Use those questions, or close versions of them, as headings, and answer directly underneath. That gives an agent a clear outline to search, and the heading itself matches the question it was asked. Our guide to <a href=\"https://ghostagentlab.com/articles/titles-descriptions-headings-ai/\">titles, descriptions and headings</a> covers how to build that outline.</p>\n          <p>A shipping page might look like this:</p>\n          <pre><code>h1  Shipping and delivery\n  h2  Where do you ship?\n  h2  How much does shipping cost?\n  h2  How long does delivery take?\n  h2  Can I track my order?\n  h2  What if my order arrives damaged?</code></pre>\n          <p>Use plain HTML tables for anything with more than two variables, such as rates by region and speed. A table reads clearly as text; a rate chart saved as an image doesn't read at all.</p>\n          <h2>Make sure agents can reach and read the page</h2>\n          <ul>\n            <li><strong>Put the text in the HTML.</strong> Accordions are fine if the answers are already in the page and just hidden until clicked. They're a problem if each answer is loaded by JavaScript when opened. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">our guide to JavaScript-only content</a>. AgentScore's JavaScript check looks at your home page, so check policy pages yourself with View Page Source.</li>\n            <li><strong>Avoid PDFs and images for policies.</strong> Some agents can read PDFs, many won't bother, and text in images is invisible to most.</li>\n            <li><strong>Don't hide policies inside a chat widget.</strong> A help bot can be useful for people, but agents need a page with a URL they can read.</li>\n            <li><strong>Link to them from every page.</strong> Footer links to Shipping, Returns, FAQ and Contact, in plain HTML, are what agents follow.</li>\n            <li><strong>List them in your sitemap and llms.txt.</strong> Your <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">sitemap</a> helps crawlers find them, and an <a href=\"https://ghostagentlab.com/articles/llms-txt/\">llms.txt file</a> can point agents straight at them with a one-line description each.</li>\n            <li><strong>Don't block them.</strong> Check that robots.txt and bot protection let AI assistants read help pages. See <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>.</li>\n          </ul>\n          <h2>Keep one version of the truth</h2>\n          <p>Contradictions are the fastest way to an inaccurate answer. If the FAQ says 30-day returns, the returns page says 14 days and a product page says \"free returns for 60 days\", an assistant has to pick one, and it may not pick the one you meant.</p>\n          <ol>\n            <li><strong>Choose a home for each policy.</strong> The returns page owns the returns policy. The FAQ summarizes it and links there.</li>\n            <li><strong>Search your site for old terms</strong> after any policy change: old return windows, old free-shipping thresholds, discontinued services.</li>\n            <li><strong>Match your structured data.</strong> If your product markup includes return or shipping details, update it with the page. See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">Product data that AI shopping agents can read</a>.</li>\n            <li><strong>Show a \"last updated\" date</strong> at the top of each policy page, so agents and people can tell how current it is.</li>\n            <li><strong>Handle temporary changes explicitly.</strong> For holiday cut-off dates or a carrier delay, add a dated note at the top (\"Orders placed after December 18 may arrive after December 25\") and remove it when it no longer applies.</li>\n          </ol>\n          <h2>Different terms for different places or products</h2>\n          <p>If your policies vary by country, membership or product type, say so clearly and early. One page per region, or a clearly labelled section for each, works better than footnotes. Name the region in the heading: \"Returns from Canada\", not \"Other regions\". If marketplace sellers on your site set their own policies, say that too, and say where the shopper can find them.</p>\n          <div>\n            <p><strong>Test it the way your customers will:</strong> ask two or three AI assistants the questions your support team hears most, naming your store. \"What's Northwind Coffee's return policy?\" \"Does Northwind ship to Canada, and how much does it cost?\" Compare the answers with your actual policy. Where they're wrong or hesitant, the page that should have answered is the one to rewrite.</p>\n          </div>\n          <h2>A short checklist</h2>\n          <ul>\n            <li>Every policy has numbers, conditions and exceptions in plain sentences.</li>\n            <li>Headings are the questions customers ask.</li>\n            <li>The text is in the HTML, not in images, PDFs or JavaScript-loaded panels.</li>\n            <li>Shipping, Returns, FAQ and Contact are linked from every page and listed in the sitemap.</li>\n            <li>No two pages disagree, and each policy page shows when it was last updated.</li>\n          </ul>\n          <p>Policy pages are where an assistant's answer turns into a sale or a lost one. Once they're clear, run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> to check the rest of your site, and read the <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness guide</a> for the bigger picture.</p>",
      "date_published": "2026-10-09T12:06:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/",
      "url": "https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/",
      "title": "Cookie banners and pop-ups that trap AI agents",
      "summary": "How cookie banners, sign-up pop-ups and chat widgets block AI agents, what AgentScore checks, and how to design overlays agents can dismiss.",
      "content_html": "<p>The first thing many AI agents see on a website isn't the home page. It's a cookie banner, a discount pop-up or a country selector sitting on top of it. A person closes it without thinking. An agent has to find the right button, work out what it means, and press it, and if it can't, the visit ends there. This guide covers the overlays that cause trouble, what AgentScore checks, and how to design them so they stay out of the way.</p>\n          <h2>Why overlays stop AI agents</h2>\n          <p>Browser agents, the kind that open a real browser and click through a site for someone, read a page through its structure and often a screenshot as well. An overlay that covers the page changes both. In the screenshot, the content is hidden or dimmed. In the page structure, a modal can make everything behind it unreachable until it's closed.</p>\n          <p>So the agent's first job is to get rid of it. That goes wrong in a few predictable ways:</p>\n          <ul>\n            <li><strong>No button it can recognize.</strong> A close control drawn as an icon with no text or label has no name, so the agent can't tell it apart from any other shape on the screen.</li>\n            <li><strong>A fake button.</strong> A styled <code>&lt;div&gt;</code> or <code>&lt;span&gt;</code> that responds to a mouse click but isn't a real button. Many agents won't recognize it as something they can press.</li>\n            <li><strong>Unclear choices.</strong> \"Customize\", \"Learn more\" and \"OK\" side by side, or a reject option hidden behind a settings screen. The agent may pick the wrong one or stall.</li>\n            <li><strong>Overlays that come back.</strong> A pop-up that reappears on every page, or a second pop-up after the first is closed, costs the agent a step every time.</li>\n            <li><strong>Something sitting on the main button.</strong> A chat widget or sticky banner placed over \"Add to cart\" means the agent's click lands on the wrong thing.</li>\n          </ul>\n          <p>For the business, this is a quiet loss. The agent doesn't complain or fill in a feedback form. It tells the person it couldn't get through, or moves on to a site where it could.</p>\n          <h2>What AgentScore checks</h2>\n          <p>The <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> check is called <strong>\"Pop-ups and banners can be dismissed by agents\"</strong>, and it's one of the higher-weighted checks in the Navigability category. It works like this:</p>\n          <ol>\n            <li>It opens your home page in a real browser at desktop size and waits for it to load.</li>\n            <li>It looks for anything fixed to the screen (a pop-up, a banner, a drawer) that covers 15% or more of the visible window. A slim header bar along the top doesn't count.</li>\n            <li>If it finds one, it reads the buttons and links inside it and looks for one with a clear dismiss label, such as \"Accept\", \"Reject\", \"Decline\", \"Close\", \"Got it\", \"No thanks\", \"Continue\" or a labelled ×.</li>\n          </ol>\n          <table>\n            <thead><tr><th>What AgentScore finds</th><th>Result</th></tr></thead>\n            <tbody>\n              <tr><td>Nothing covers the page on arrival</td><td>Pass</td></tr>\n              <tr><td>An overlay covers the page, but it has a clearly labelled button to close or answer it</td><td>Warning: agents can get past it, but every one has to deal with it first</td></tr>\n              <tr><td>An overlay covers the page and has no clearly labelled way out</td><td>Fail</td></tr>\n            </tbody>\n          </table>\n          <p>A second check, <strong>\"Agents can find and press your add-to-cart or sign-up button\"</strong>, catches the other half of the problem. It finds the main button on a product or pricing page and tests whether something else, usually a pop-up, banner or chat widget, sits on top of it. If it does, the check fails, because an agent's click would land on the overlay instead.</p>\n          <div>\n            <p><strong>This isn't about skipping consent.</strong> Where the law requires you to ask for cookie consent, keep asking. The goal is a banner that an AI agent acting for a person can answer as easily as the person could, not one that disappears. Talk to whoever owns privacy before changing what your banner asks or when.</p>\n          </div>\n          <h2>The usual suspects</h2>\n          <table>\n            <thead><tr><th>Overlay</th><th>What tends to go wrong</th><th>Better approach</th></tr></thead>\n            <tbody>\n              <tr><td>Cookie consent banner</td><td>Full-screen modal; icon-only close; reject hidden in settings</td><td>A bar at the bottom of the screen with \"Accept all\" and \"Reject all\" as real, named buttons side by side</td></tr>\n              <tr><td>Newsletter or discount pop-up</td><td>Appears on arrival, covers the page, close button is a small unlabelled icon</td><td>Wait until someone has engaged; use a labelled \"Close\" or \"No thanks\"; don't show it on product or checkout pages</td></tr>\n              <tr><td>Country or region selector</td><td>Blocks the page until a choice is made, often with a custom dropdown</td><td>A small, dismissible suggestion (\"Shopping from Canada? Go to the Canadian store\") and locale in the URL</td></tr>\n              <tr><td>Age check</td><td>Custom date pickers or a fake button</td><td>Where required, a short form with a native field and a real \"Confirm\" button</td></tr>\n              <tr><td>Chat widget</td><td>Opens itself on load, or its launcher covers buttons on smaller screens</td><td>Start collapsed, keep the launcher small and clear of the main buttons</td></tr>\n              <tr><td>App install or promo banner</td><td>Sticky and tall, pushing content down or covering it</td><td>Keep it short, dismissible with a named button, and remember the choice</td></tr>\n            </tbody>\n          </table>\n          <h2>What an agent-friendly banner looks like</h2>\n          <p>Here's a cookie banner that people, screen readers and AI agents can all handle. It sits at the bottom of the page, doesn't block what's behind it, and every choice is a real button with a plain name.</p>\n          <figure>\n            <div>\n              <div>\n                <p>Traps agents</p>\n                <ul>\n                  <li>Covers the whole page until it&rsquo;s answered</li>\n                  <li>The &times; has no name an agent can read</li>\n                  <li>Rejecting is hidden behind &ldquo;Settings&rdquo;</li>\n                </ul>\n              </div>\n              <div>\n                <p>Works for agents</p>\n                <ul>\n                  <li>Sits at the edge, so the page stays usable</li>\n                  <li>Every choice is a real button with words</li>\n                  <li>Rejecting takes one step, like accepting</li>\n                </ul>\n              </div>\n            </div>\n            <figcaption>Two cookie banners. The first stops every agent until it&rsquo;s answered and gives it nothing it can name. The second can be answered in one step, or left alone.</figcaption>\n          </figure>\n          <pre><code>&lt;div class=\"consent-bar\" role=\"region\" aria-label=\"Cookie choices\"&gt;\n  &lt;p&gt;We use cookies to run the store and, with your OK, to measure ads.\n     &lt;a href=\"/pages/cookies\"&gt;Cookie policy&lt;/a&gt;&lt;/p&gt;\n  &lt;button type=\"button\"&gt;Reject all&lt;/button&gt;\n  &lt;button type=\"button\"&gt;Accept all&lt;/button&gt;\n  &lt;button type=\"button\"&gt;Choose cookies&lt;/button&gt;\n&lt;/div&gt;</code></pre>\n          <p>And the most common failure, the icon-only close button, fixed in one attribute:</p>\n          <pre><code>&lt;!-- An agent sees: button, no name --&gt;\n&lt;button class=\"modal-x\"&gt;&lt;svg&gt;…&lt;/svg&gt;&lt;/button&gt;\n&lt;!-- An agent sees: button \"Close\" --&gt;\n&lt;button class=\"modal-x\" aria-label=\"Close\"&gt;&lt;svg aria-hidden=\"true\"&gt;…&lt;/svg&gt;&lt;/button&gt;</code></pre>\n          <p>A few rules that make the difference:</p>\n          <ul>\n            <li><strong>Use real buttons with words.</strong> \"Accept all\", \"Reject all\", \"Close\", \"No thanks\". If the design calls for an icon, give it an <code>aria-label</code>.</li>\n            <li><strong>Keep it to the edge.</strong> A bar along the bottom leaves the page usable. A centered modal over a dimmed page forces every visitor, human or agent, to deal with it first.</li>\n            <li><strong>Make rejecting as easy as accepting.</strong> Several European privacy regulators have said this about cookie banners, and it also means an agent told to decline tracking can do so in one step.</li>\n            <li><strong>Let Escape close it</strong> where closing is allowed, and return focus to the page afterwards.</li>\n            <li><strong>Remember the answer.</strong> Once a choice is made, don't ask again on the next page.</li>\n            <li><strong>One overlay at a time.</strong> Don't stack a discount pop-up on top of a cookie banner on the first page view.</li>\n          </ul>\n          <h2>Pop-ups and the pages that earn money</h2>\n          <p>The home page is where AgentScore looks for overlays, but the pages that matter most are product pages, the cart and checkout. That's where a covered button costs a sale. Check those pages on a narrower screen too: a chat launcher that sits harmlessly in a corner on a laptop can land on top of \"Add to cart\" on a phone-sized window.</p>\n          <p>If marketing relies on an email capture pop-up, that's a reasonable trade-off to discuss rather than a rule to break. Showing it after someone has scrolled or spent time on the page, rather than on arrival, keeps most of its value for people and removes it from the first moment an agent arrives.</p>\n          <h2>How to test it</h2>\n          <ol>\n            <li><strong>Open the site in a private window.</strong> That's what a first-time visitor and most agents see: no cookies, no remembered choices.</li>\n            <li><strong>Try to get past every overlay with the keyboard alone.</strong> Tab to the buttons and press Enter. If you can't reach the close button, or can't tell which control has focus, an agent will struggle too.</li>\n            <li><strong>Check the names.</strong> In Chrome's developer tools, the Accessibility pane shows each button's name. An empty name is a problem.</li>\n            <li><strong>Visit a product page and the cart</strong> at a narrow window size and make sure nothing covers the main button.</li>\n            <li><strong>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>.</strong> It reports any overlay it found, how much of the screen it covered and which buttons it offered, and checks whether your add-to-cart or sign-up button is covered.</li>\n          </ol>\n          <p>Overlays are one part of how agents find their way around a page. For the rest, see <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>, and for the other thing that often greets agents on arrival, a CAPTCHA or bot challenge, see <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a>. To see how agents handle the whole journey, not just the first screen, read <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">how to test your key journeys with AI agents</a>.</p>",
      "date_published": "2026-10-09T12:05:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Navigability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/link-key-pages-home-page/",
      "url": "https://ghostagentlab.com/articles/link-key-pages-home-page/",
      "title": "Link your key pages from the home page",
      "summary": "AI agents start at your home page and follow links. Make sure products, the cart and pricing are linked in plain HTML so every agent can find them.",
      "content_html": "<p>When an AI agent is asked about your business, it usually starts where people do: your home page. From there it follows links. If the links to your products, cart or pricing aren't in the page's HTML, many agents never find those pages, however good they are. It's one of the simplest things to get right, and one of the easiest to break without noticing.</p>\n          <h2>How agents find their way from the home page</h2>\n          <p>An AI assistant answering \"how much is the Pro plan at northwind.example?\" or an agent asked to \"buy two bags of the house blend\" needs to get from your home page to the right page quickly. Agents differ in how they do that:</p>\n          <ul>\n            <li><strong>Agents that read the HTML.</strong> Many AI assistants and AI search crawlers fetch the page as your server sends it and read the links in it. They don't run JavaScript, so a menu that's built in the browser is invisible to them. (See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see your JavaScript-only content</a>.)</li>\n            <li><strong>Browser agents.</strong> These open a real browser and click around. They can see menus built with JavaScript, but they still need real links and clear names to choose from, and every extra click is a chance to get lost.</li>\n          </ul>\n          <p>Either way, the home page is the map. If the pages that make you money aren't on it, the agent has to guess, search, or fall back on your sitemap, if it looks for one at all.</p>\n          <h2>What AgentScore checks</h2>\n          <p><a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> goes looking for your key pages the way an agent would, in this order:</p>\n          <ol>\n            <li><strong>Links in the home page's HTML.</strong> It reads every link to your own site in the page as your server sends it, before any JavaScript runs.</li>\n            <li><strong>Your sitemap,</strong> if nothing suitable was linked.</li>\n            <li><strong>A couple of common addresses,</strong> such as <code>/pricing</code> or <code>/cart</code>, for anything still missing.</li>\n          </ol>\n          <p>It recognizes a key page by its address or its link text. Addresses like <code>/products/house-blend</code>, <code>/pricing</code>, <code>/plans</code> and <code>/cart</code> count, and so do links labelled \"Pricing\", \"Plans\", \"Cart\", \"Basket\" or \"Bag\".</p>\n          <p>The check, <strong>\"Key pages are linked from the home page\"</strong>, then asks one question: did the home page link to them directly?</p>\n          <table>\n            <thead><tr><th>Type of site</th><th>Passes when the home page HTML links to</th></tr></thead>\n            <tbody>\n              <tr><td>Online store</td><td>A product page or the cart</td></tr>\n              <tr><td>Software or subscription business</td><td>The pricing page</td></tr>\n              <tr><td>Neither found anywhere</td><td>Not tested</td></tr>\n            </tbody>\n          </table>\n          <p>If AgentScore only found your key pages through the sitemap or a common address, the check gives a warning: the pages exist, but an agent reading your home page wouldn't have been led to them. The same pages are then used for other checks, such as whether AI agents can open them and whether prices are readable, so linking them clearly helps more than one result.</p>\n          <h2>Why links go missing</h2>\n          <p>Most sites that get a warning here do have links to their products and pricing. People can see them. The links just aren't where agents look.</p>\n          <ul>\n            <li><strong>Menus built by JavaScript.</strong> A mega menu or mobile menu that's fetched and drawn after the page loads isn't in the HTML at all.</li>\n            <li><strong>Buttons pretending to be links.</strong> A \"Shop now\" that's a <code>&lt;button&gt;</code> or <code>&lt;div&gt;</code> with a click handler goes nowhere for an agent that reads links. So does <code>href=\"#\"</code> with a script.</li>\n            <li><strong>Vague link text and addresses.</strong> \"Discover\", \"Explore the range\" or \"See what it costs\", pointing to <code>/landing/q4-campaign</code>, gives neither people nor agents much to go on.</li>\n            <li><strong>Product links only in a carousel</strong> that's loaded after the page renders, or only inside a search box.</li>\n            <li><strong>A cart icon with no link.</strong> A cart drawer that opens with a script, with no link to a cart page behind it.</li>\n          </ul>\n          <h2>What a good home page links to</h2>\n          <p>Think of the questions people ask AI assistants about businesses like yours, and make sure the answer is one link away from the home page.</p>\n          <table>\n            <thead><tr><th>Type of site</th><th>Link to from the home page</th></tr></thead>\n            <tbody>\n              <tr><td>Online store</td><td>Main categories, a few bestselling products, the cart, shipping, returns, contact</td></tr>\n              <tr><td>Software company</td><td>Pricing, sign-up, product overview, docs, contact sales</td></tr>\n              <tr><td>Service business</td><td>Services and prices, booking, locations and hours, contact</td></tr>\n            </tbody>\n          </table>\n          <p>Footer links count. A plain footer with \"Pricing\", \"Shipping\" and \"Returns\" is often the most reliable route an agent has, because it's on every page and rarely depends on JavaScript.</p>\n          <h2>What the HTML should look like</h2>\n          <p>The fix is ordinary links, present in the HTML your server sends, with clear text:</p>\n          <pre><code>&lt;!-- Hard for agents: no link in the HTML, vague text --&gt;\n&lt;div class=\"nav-item\" onclick=\"openMenu('shop')\"&gt;Discover&lt;/div&gt;\n&lt;span class=\"cart-icon\" onclick=\"openCart()\"&gt;&lt;/span&gt;\n&lt;!-- Easy for agents: real links, plain names --&gt;\n&lt;nav aria-label=\"Main\"&gt;\n  &lt;a href=\"/collections/coffee\"&gt;Coffee&lt;/a&gt;\n  &lt;a href=\"/collections/equipment\"&gt;Brewing equipment&lt;/a&gt;\n  &lt;a href=\"/pricing\"&gt;Pricing&lt;/a&gt;\n  &lt;a href=\"/cart\" aria-label=\"Cart\"&gt;&lt;svg aria-hidden=\"true\"&gt;…&lt;/svg&gt;&lt;/a&gt;\n&lt;/nav&gt;</code></pre>\n          <ul>\n            <li><strong>Use <code>&lt;a href&gt;</code> with a real address</strong> for anything that goes to another page. A cart drawer is fine, as long as the cart icon is also a link to a cart page.</li>\n            <li><strong>Render the main menu on the server.</strong> If your platform or theme builds the menu in the browser, ask your developer to output at least the top-level links in the HTML.</li>\n            <li><strong>Name links for what's behind them.</strong> \"Pricing\" beats \"See what it costs\"; \"Shop coffee\" beats \"Discover\".</li>\n            <li><strong>Keep addresses readable.</strong> <code>/pricing</code>, <code>/cart</code> and <code>/products/house-blend</code> tell an agent what a page is before it opens it.</li>\n            <li><strong>Make menus open on click, not only on hover,</strong> and make the top-level items real links too, so there's always a route in.</li>\n          </ul>\n          <h2>How to check your own home page</h2>\n          <ol>\n            <li><strong>View the source.</strong> In most browsers, right-click the page and choose \"View page source\" (not \"Inspect\", which shows the page after JavaScript has run). Search for <code>href=\"/cart\"</code>, <code>/pricing</code> or <code>/products/</code>. If they're not there, agents that read HTML can't see them.</li>\n            <li><strong>Turn JavaScript off</strong> in your browser's settings and reload the home page. Whatever navigation is left is what those agents get.</li>\n            <li><strong>Click through with the keyboard.</strong> Tab from the top of the page to a product and the cart. If you can't, browser agents may not either.</li>\n            <li><strong>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>.</strong> It tells you whether your products, cart or pricing were linked from the home page, or only turned up another way, such as through the sitemap.</li>\n          </ol>\n          <div>\n            <p><strong>Your sitemap is a safety net, not a substitute.</strong> A good <a href=\"https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/\">XML sitemap</a> helps crawlers find every page, and <a href=\"https://ghostagentlab.com/articles/llms-txt/\">llms.txt</a> can point AI straight to your most important ones. But plenty of agents never read either. Links on the home page are the route every agent can follow.</p>\n          </div>\n          <p>Once agents can reach your key pages, the next questions are whether they can read them and act on them. See <a href=\"https://ghostagentlab.com/articles/machine-readable-prices/\">prices AI agents can read</a>, <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a> and <a href=\"https://ghostagentlab.com/articles/site-search-filters-ai-agents/\">site search and filters AI agents can use</a>.</p>",
      "date_published": "2026-10-09T12:04:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Navigability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/site-search-filters-ai-agents/",
      "url": "https://ghostagentlab.com/articles/site-search-filters-ai-agents/",
      "title": "Site search and filters AI agents can use",
      "summary": "Build a search box, results URLs and filters that AI agents can operate, so they find the right products instead of giving up or guessing.",
      "content_html": "<p>Ask an AI agent to \"find a waterproof trail shoe under $150 in size 10\" and, on most stores, it does what a person would: types into the search box, then narrows the results with filters. If the search box is hard to find, the results only exist inside a script, or the filters are custom widgets it can't operate, the agent gives up or picks the wrong thing. This guide covers how to build search and filters that AI agents can use, and that work better for people too.</p>\n          <h2>How agents use search and filters</h2>\n          <p>There are two broad ways an AI agent gets to a list of products:</p>\n          <ul>\n            <li><strong>Through the page.</strong> Browser agents open your site, find the search box by its label, type, press Enter, and then click filters and sort options, much as a keyboard or screen-reader user would.</li>\n            <li><strong>Through the address.</strong> If your search and filtered results have their own URLs, an agent can go straight to <code>/search?q=trail+shoes</code> or a filtered category page. AI assistants that fetch pages rather than click through them can only work this way: they can't type into a box, but they can open a link.</li>\n          </ul>\n          <p>A site that supports both gives every kind of agent a route to the right products. A site that supports neither forces agents to browse category pages one by one, which is slow and often ends with a worse answer.</p>\n          <h2>A search box agents can find and use</h2>\n          <p>The search box should be a real form that works without anything clever. That means a labelled field, a named submit button, and a results page that loads when the form is submitted.</p>\n          <pre><code>&lt;form role=\"search\" action=\"/search\" method=\"get\"&gt;\n  &lt;label for=\"q\"&gt;Search products&lt;/label&gt;\n  &lt;input id=\"q\" name=\"q\" type=\"search\" autocomplete=\"off\"&gt;\n  &lt;button type=\"submit\"&gt;Search&lt;/button&gt;\n&lt;/form&gt;</code></pre>\n          <ul>\n            <li><strong>Label it.</strong> A visible label or an <code>aria-label</code> such as \"Search products\". A magnifying-glass icon on its own has no name an agent can read.</li>\n            <li><strong>Keep it visible.</strong> A search box that only appears after clicking an unlabelled icon is an extra step agents often miss. If space is tight, make the icon a real button named \"Search\".</li>\n            <li><strong>Make Enter work.</strong> Pressing Enter should submit the search and open a results page. Predictive suggestions as you type are fine, but they shouldn't be the only way to get results.</li>\n            <li><strong>Accept the words people use.</strong> Product names, product types, brands, SKUs, common misspellings and synonyms (\"sneakers\" and \"trainers\"). An agent passes on what the person asked for, in their words.</li>\n          </ul>\n          <h2>Give results and filters their own URLs</h2>\n          <p>This is the single most useful change for AI agents, and it's good for people too: results can be bookmarked, shared and reached with the back button.</p>\n          <p>Every search and every combination of filters should be reflected in the address, using readable parameters:</p>\n          <pre><code>https://northwind.example/search?q=trail+shoes\nhttps://northwind.example/collections/shoes?waterproof=yes&amp;size=10&amp;max_price=150\nhttps://northwind.example/collections/shoes?sort=price-asc&amp;page=2</code></pre>\n          <ul>\n            <li><strong>Update the URL when filters change.</strong> If your filters work without reloading the page, update the address as they're applied, so the current view always has a link.</li>\n            <li><strong>Make the URL load the same results.</strong> Opening that link fresh, with no history, should show the same filtered list.</li>\n            <li><strong>Put results in the HTML.</strong> If results are only drawn by JavaScript, agents that read HTML see an empty page. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see your JavaScript-only content</a>.</li>\n            <li><strong>Use plain parameter names and values.</strong> <code>size=10</code> and <code>color=black</code> are guessable. <code>f=a8x2k</code> isn't.</li>\n          </ul>\n          <h2>Filters as real controls</h2>\n          <p>Filters are where many stores lose agents. They're often built as custom widgets that look good but aren't recognizable as checkboxes, options or buttons. The fix is to use the controls the browser already provides.</p>\n          <table>\n            <thead><tr><th>Filter</th><th>Hard for agents</th><th>Easy for agents</th></tr></thead>\n            <tbody>\n              <tr><td>Size, brand, feature</td><td>Clickable boxes made from <code>&lt;div&gt;</code>s</td><td>Checkboxes with labels, grouped in a <code>&lt;fieldset&gt;</code> with a <code>&lt;legend&gt;</code></td></tr>\n              <tr><td>Color</td><td>Swatches with no text, only a background color</td><td>Checkboxes or buttons named \"Black\", \"Navy\", with the swatch as decoration</td></tr>\n              <tr><td>Price</td><td>A drag-only slider</td><td>\"Min\" and \"Max\" number fields, with or without a slider alongside</td></tr>\n              <tr><td>Sort order</td><td>A custom dropdown that opens on hover</td><td>A native <code>&lt;select&gt;</code>, or a list of real links</td></tr>\n              <tr><td>More results</td><td>Infinite scroll only</td><td>Numbered page links, or a \"Load more\" button backed by a <code>?page=2</code> link</td></tr>\n            </tbody>\n          </table>\n          <pre><code>&lt;fieldset&gt;\n  &lt;legend&gt;Size&lt;/legend&gt;\n  &lt;label&gt;&lt;input type=\"checkbox\" name=\"size\" value=\"9\"&gt; 9 (14)&lt;/label&gt;\n  &lt;label&gt;&lt;input type=\"checkbox\" name=\"size\" value=\"10\"&gt; 10 (22)&lt;/label&gt;\n&lt;/fieldset&gt;\n&lt;label for=\"sort\"&gt;Sort by&lt;/label&gt;\n&lt;select id=\"sort\" name=\"sort\"&gt;\n  &lt;option value=\"relevance\"&gt;Most relevant&lt;/option&gt;\n  &lt;option value=\"price-asc\"&gt;Price: low to high&lt;/option&gt;\n&lt;/select&gt;</code></pre>\n          <p>A few more details that save agents from guessing:</p>\n          <ul>\n            <li><strong>Say how many results there are</strong> in text (\"24 products\"), and update it when filters change. An <code>aria-live</code> region lets assistive tech, and agents reading the page structure, notice the change.</li>\n            <li><strong>Show active filters</strong> as named buttons that remove them (\"Remove filter: Size 10\"), plus a \"Clear all\".</li>\n            <li><strong>Handle no results helpfully.</strong> Say plainly that nothing matched, and offer a way out: related categories, or the same search without the last filter.</li>\n            <li><strong>Keep filter panels reachable.</strong> On smaller screens, the \"Filters\" toggle should be a real, named button, not an icon.</li>\n          </ul>\n          <h2>Tell agents how your search works</h2>\n          <p>Once your search has a predictable URL, you can describe it so agents don't have to discover it:</p>\n          <ul>\n            <li><strong>In llms.txt.</strong> A line such as \"Search: <code>https://northwind.example/search?q={query}</code>\" in your <a href=\"https://ghostagentlab.com/articles/llms-txt/\">llms.txt file</a> hands the pattern to any AI that reads it.</li>\n            <li><strong>In structured data.</strong> schema.org lets a <code>WebSite</code> describe its search with a <a href=\"https://schema.org/SearchAction\">SearchAction</a>. Google retired the search box it used to show for this in 2024, so don't expect a search result feature from it, but it remains a standard, machine-readable description of your search URL.</li>\n            <li><strong>As a tool.</strong> The most direct option is to let agents call your search as a tool, through an MCP server or WebMCP in the browser. Both are newer and still settling; see <a href=\"https://ghostagentlab.com/articles/mcp-webmcp-for-websites/\">MCP and WebMCP for websites</a>.</li>\n          </ul>\n          <div>\n            <p><strong>Watch your robots.txt.</strong> Many sites stop search engines indexing internal search results, which is sensible. But a blanket <code>Disallow: /search</code> for every bot also tells AI assistants that respect robots.txt not to open those pages when a person asks them to. Decide deliberately which agents that rule should apply to. See <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>.</p>\n          </div>\n          <h2>What AgentScore and Ghost Agents test</h2>\n          <p>AgentScore doesn't have a separate check for search, but several of its checks cover the building blocks. On the home page, where most search boxes live, it checks that form fields are labelled, that buttons and links have names, and that clickable things are real buttons and links. \"Content loads without JavaScript\" catches a home page whose content only appears once scripts run, a sign that results and filters may be built the same way. The free scan doesn't run a shopping journey itself; a <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> test does, and a search box or filter that doesn't work shows up in its step-by-step replay.</p>\n          <p>Search is also one of the journeys <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> can test for you. You choose the journey, such as \"search for a waterproof trail shoe, filter to size 10 and open the first result\", and a Ghost Agent works through it like a customer would, on the schedule you set. When a theme update or a new search app breaks it, you find out before customers do. See <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">how to test your key journeys with AI agents</a>.</p>\n          <h2>A quick checklist</h2>\n          <ol>\n            <li>Search box is visible, labelled, and submits with Enter to a results page.</li>\n            <li>Searches and filtered views have their own readable URLs that load the same results when opened fresh.</li>\n            <li>Results are in the page HTML, not only drawn by JavaScript.</li>\n            <li>Filters are checkboxes, selects, number fields and buttons with names, not styled boxes.</li>\n            <li>Result counts, active filters and \"no results\" are shown in text.</li>\n            <li>Pagination uses real links.</li>\n            <li>Your search URL is described in llms.txt, and robots.txt doesn't shut AI assistants out of it by accident.</li>\n          </ol>\n          <p>For the controls on the rest of the page, see <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>. For getting agents to your categories in the first place, see <a href=\"https://ghostagentlab.com/articles/link-key-pages-home-page/\">link your key pages from the home page</a>.</p>",
      "date_published": "2026-10-09T12:03:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Navigability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/agent-friendly-sign-up-booking/",
      "url": "https://ghostagentlab.com/articles/agent-friendly-sign-up-booking/",
      "title": "Sign-up, booking and lead forms AI agents can finish",
      "summary": "How to build sign-up, booking and quote forms AI agents can complete, and how to handle email verification, CAPTCHAs, date pickers and the handoff.",
      "content_html": "<p>\"Book me a table for four on Friday.\" \"Get me a quote for a new roof.\" \"Sign me up for the free trial.\" More of these requests now go to an AI assistant instead of a search box, and the assistant has to finish the form on your site. If it can't, the lead goes to a competitor whose form it can. This guide covers sign-up, booking and lead forms, and the steps where agents most often get stuck.</p>\n          <h2>Why forms are where agents give up</h2>\n          <p>Reading a page is easy for an AI agent. Finishing a form is not. Every field has to be matched to something the person asked for, every choice has to be made, and every step has to confirm that it worked. Forms also tend to carry the things built to stop bots: CAPTCHAs, email verification and one-time codes.</p>\n          <p>None of this means removing your defenses. It means designing the form so an agent can do the routine part, and handing over to the person at the moments that genuinely need them.</p>\n          <h2>The basics every form needs</h2>\n          <ul>\n            <li><strong>A visible label on every field.</strong> Agents match \"my phone number\" to a field called \"Phone number\". Placeholder text alone is a weak substitute. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>.</li>\n            <li><strong>The right field types and autocomplete values.</strong> <code>type=\"email\"</code>, <code>type=\"tel\"</code>, <code>autocomplete=\"given-name\"</code>, <code>postal-code</code> and <code>organization</code> tell an agent exactly what each field wants.</li>\n            <li><strong>Only the fields you need.</strong> Every optional field is another chance to guess wrong. Ask for the rest after the lead or booking is in.</li>\n            <li><strong>Native controls for choices.</strong> Real <code>&lt;select&gt;</code> menus, radio buttons and checkboxes, not custom dropdowns built from <code>&lt;div&gt;</code>s.</li>\n            <li><strong>Errors written as text next to the field.</strong> \"Enter a phone number with area code\" tells the agent what to fix. A red border doesn't.</li>\n            <li><strong>A clear confirmation.</strong> After submitting, show a page or message that says what happened: \"Thanks, Sam. Your table for 4 is booked for Friday 10 October at 7:30 pm. Reference NW-4821.\" The agent can read that back to the person. A toast that disappears after two seconds can't be.</li>\n          </ul>\n          <h2>Sign-up forms</h2>\n          <p>For software companies, sign-up is the conversion. An agent asked to \"start a trial\" needs to find the sign-up page, fill it in and reach a confirmed account, or at least a clear next step for the person.</p>\n          <ul>\n            <li><strong>Link to sign-up from every page.</strong> A real link with clear text (\"Start free trial\"), not a button that opens a modal only after a script loads. AgentScore checks that agents can find and press your sign-up button.</li>\n            <li><strong>Keep password rules on the page.</strong> State them before the field (\"At least 12 characters\"), not only in an error after submit.</li>\n            <li><strong>Make \"Sign up with Google\" an option, not the only route.</strong> Third-party sign-in opens a pop-up on another domain that many agents can't or shouldn't use. An email option keeps the door open.</li>\n            <li><strong>Don't hide sign-up behind a sales call.</strong> If your plans need a conversation, say so in text, and offer a booking form the agent can complete.</li>\n          </ul>\n          <h2>Email verification and one-time codes</h2>\n          <p>An agent working on someone's behalf usually can't read their inbox or their text messages, and you wouldn't want it to have that access by default. So any step that says \"enter the code we just sent\" is a handoff point. Handle it well:</p>\n          <ol>\n            <li><strong>Let people do as much as possible before verification.</strong> Create the account, save the booking request or record the lead first, then ask for confirmation. If the agent stops at the verification step, the work isn't lost.</li>\n            <li><strong>Say clearly what happens next.</strong> \"We've sent a link to sam@example.com. Click it to activate your account.\" The agent can pass that message straight to the person.</li>\n            <li><strong>Prefer a link over a short-lived code.</strong> A link the person clicks later works across devices and doesn't expire while the agent reports back. If you use codes, give them a sensible lifetime and a \"resend\" button with a name.</li>\n            <li><strong>Mark code fields properly.</strong> <code>autocomplete=\"one-time-code\"</code> and <code>inputmode=\"numeric\"</code> on a single field are easier for people and agents than six separate boxes.</li>\n          </ol>\n          <h2>CAPTCHAs on forms</h2>\n          <p>A CAPTCHA on every submission stops AI agents as surely as it stops spam bots. Most sites don't need one on every form. Better options:</p>\n          <ul>\n            <li><strong>Risk-based challenges</strong> that only appear for suspicious submissions, such as many attempts from one address in a short time.</li>\n            <li><strong>Invisible checks</strong> like honeypot fields and rate limits, which stop crude bots without asking anyone to solve a puzzle.</li>\n            <li><strong>Recognizing verified agents.</strong> Some agents sign their requests with <a href=\"https://ghostagentlab.com/blog/signed-requests/\">Web Bot Auth</a>, and services such as Cloudflare can let verified bots through.</li>\n          </ul>\n          <p>AgentScore checks for a CAPTCHA or challenge when an agent arrives on your site, and at the cart and checkout. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a> for the full picture.</p>\n          <h2>Booking forms and date pickers</h2>\n          <p>Booking is where custom widgets cause the most trouble. A calendar that only works by clicking tiny cells, or a time-slot grid of unlabelled boxes, leaves an agent guessing.</p>\n          <table>\n            <thead><tr><th>Element</th><th>Hard for agents</th><th>Easy for agents</th></tr></thead>\n            <tbody>\n              <tr><td>Date</td><td>Custom calendar pop-up with no labels on days</td><td><code>&lt;input type=\"date\"&gt;</code>, or a calendar whose days are buttons named \"Friday 10 October, available\"</td></tr>\n              <tr><td>Time</td><td>Grid of colored boxes</td><td>Radio buttons or a list of buttons named \"7:30 pm\"</td></tr>\n              <tr><td>Availability</td><td>Grayed-out slots with no text</td><td>Disabled options with a reason: \"7:00 pm, fully booked\"</td></tr>\n              <tr><td>Party size or service</td><td>Plus and minus icons only</td><td>A <code>&lt;select&gt;</code> or number field with a label</td></tr>\n            </tbody>\n          </table>\n          <p>Also make availability readable as text where you can, for example a list of open times on the page. An assistant answering \"is there anything Friday evening?\" can then answer without driving the widget at all.</p>\n          <pre><code>&lt;fieldset&gt;\n  &lt;legend&gt;Time on Friday 10 October&lt;/legend&gt;\n  &lt;label&gt;&lt;input type=\"radio\" name=\"time\" value=\"19:00\" disabled&gt; 7:00 pm (fully booked)&lt;/label&gt;\n  &lt;label&gt;&lt;input type=\"radio\" name=\"time\" value=\"19:30\"&gt; 7:30 pm&lt;/label&gt;\n  &lt;label&gt;&lt;input type=\"radio\" name=\"time\" value=\"20:00\"&gt; 8:00 pm&lt;/label&gt;\n&lt;/fieldset&gt;</code></pre>\n          <p>If your booking runs inside a third-party widget in an iframe, test it. Some booking providers build accessible widgets and some don't, and you inherit whatever they ship.</p>\n          <h2>Quote and lead forms</h2>\n          <ul>\n            <li><strong>Split long forms into clear steps</strong> with a heading per step and a named \"Next\" button. Show where the agent is: \"Step 2 of 3: Your project\".</li>\n            <li><strong>Avoid free-text-only questions</strong> where a set of options would do. \"Roof type\" as radio buttons is easier to answer correctly than a blank box.</li>\n            <li><strong>Make consent explicit.</strong> Marketing opt-ins should be unticked checkboxes with clear labels. An agent should only tick them if the person asked it to.</li>\n            <li><strong>Give a real response time.</strong> \"We'll reply within one working day\" is something the agent can tell the person. \"Someone will be in touch\" isn't much help.</li>\n          </ul>\n          <h2>Plan the handoff to the person</h2>\n          <p>Some steps belong to the person: paying a deposit, accepting terms that bind them, confirming an email address, or signing in with a password only they know. Agents generally stop and hand back at these points, and that's how it should be. Your job is to make the handoff smooth:</p>\n          <ul>\n            <li>Keep what the agent entered. A booking or form that survives a pause, ideally at its own URL, lets the person pick up where the agent left off.</li>\n            <li>Put the handoff step on its own page with a clear heading, such as \"Confirm your booking\", so the agent knows it has done its part.</li>\n            <li>Show a summary of everything entered before the final step, so the person can check it quickly.</li>\n          </ul>\n          <div>\n            <p><strong>A quick test:</strong> ask an AI assistant with browsing to complete your main form using test details, and watch where it hesitates. Then try the same journey with only the keyboard. The two usually get stuck in the same places. For a structured approach, see <a href=\"https://ghostagentlab.com/articles/test-journeys-with-ai-agents/\">how to test your key journeys with AI agents</a>.</p>\n          </div>\n          <h2>How Ghost Agent Labs helps</h2>\n          <p><a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> checks the parts of this you can see from the outside: form fields are labelled, buttons and links have names, pop-ups can be dismissed, there's no CAPTCHA on arrival, and agents can find your sign-up or add-to-cart button. <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> go further: they run the journey itself, such as sign-up or a contact form, with test data you supply, on a schedule, and show a step-by-step replay when it breaks. They don't submit payments, and the sign-up template stops short of email confirmation.</p>\n          <p>Forms are only half of task completion. For stores, the other half is <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">checkout</a> and <a href=\"https://ghostagentlab.com/articles/guest-checkout-ai-agents/\">guest checkout</a>.</p>",
      "date_published": "2026-10-09T12:02:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Task completion"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/test-journeys-with-ai-agents/",
      "url": "https://ghostagentlab.com/articles/test-journeys-with-ai-agents/",
      "title": "How to test your key journeys with AI agents",
      "summary": "Which journeys to test with AI agents, how to test by hand and with Ghost Agents, what to measure, and how to test safely without real purchases.",
      "content_html": "<p>Technical checks tell you whether AI agents should be able to use your site. The only way to know whether they can is to send one through and watch. This guide covers which journeys to test, how to test them by hand with the assistants your customers already use, how to automate it, and how to keep testing safe.</p>\n          <h2>Why checks aren't enough</h2>\n          <p>A site can pass every check for labels, named buttons and structured data and still lose agents at a custom size picker, a cookie banner that reappears on the cart page, or a checkout step that needs an account. Those failures only show up when an agent actually tries to finish the job. That's why task completion carries the most weight in <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness</a>, and why it can only be measured by running real agents.</p>\n          <h2>Step 1: Pick the journeys that make money</h2>\n          <p>Don't test everything. Start with the two or three journeys where an agent failure costs you a customer:</p>\n          <table>\n            <thead><tr><th>Type of site</th><th>Journeys to test first</th></tr></thead>\n            <tbody>\n              <tr><td>Online store</td><td>Search for a product, choose options and add to cart, check out as a guest up to payment</td></tr>\n              <tr><td>Software company</td><td>Find pricing, start a trial or sign up, book a demo</td></tr>\n              <tr><td>Service business</td><td>Find prices and availability, book an appointment, request a quote</td></tr>\n              <tr><td>Any site</td><td>Find a policy (returns, shipping, cancellation), use the contact form</td></tr>\n            </tbody>\n          </table>\n          <p>Write each one the way a customer would ask an assistant: \"Find a waterproof jacket in medium under $150 and add it to the cart.\" Specific goals give clearer results than \"buy something\".</p>\n          <h2>Step 2: Test by hand with consumer assistants</h2>\n          <p>The quickest start costs nothing. Several consumer AI assistants can now browse and act on websites, usually in an \"agent\" or browsing mode. Use the ones your customers are likely to use.</p>\n          <ol>\n            <li><strong>Give it the goal and your site.</strong> \"On northwind.example, find a 12-cup coffee maker and add it to the cart. Don't buy anything.\"</li>\n            <li><strong>Watch, don't help.</strong> Most agent modes show what the agent is doing. Note every hesitation, wrong click and retry, not just whether it finished.</li>\n            <li><strong>Ask it what went wrong.</strong> If it fails, ask what it was trying to do and what it couldn't find. The answer often names the exact control.</li>\n            <li><strong>Repeat it.</strong> Agents don't behave the same way every time. Run each journey at least three times before drawing conclusions.</li>\n          </ol>\n          <p>Manual testing is good for discovery: it shows you how agents experience your site and builds a shared picture for the team. It doesn't scale, and it won't tell you when something breaks next Tuesday.</p>\n          <h2>Step 3: Automate it with Ghost Agents</h2>\n          <p><a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> are AI agents that run your key journeys for you. In the Ghost Agent Labs app, you set up a test for a site you've added:</p>\n          <ul>\n            <li><strong>Goal.</strong> Start from a common journey (checkout, product search, sign-up or contact form) or describe your own in plain words.</li>\n            <li><strong>Start page.</strong> Where the agent begins. It can only visit pages on your site.</li>\n            <li><strong>When it passes.</strong> A check on the final page the agent reaches: it reached the payment step, the page shows certain text, or the address contains a certain path. These are checked directly, not by asking an AI model whether it succeeded.</li>\n            <li><strong>Test data.</strong> A test email address and any other values the journey needs.</li>\n            <li><strong>Schedule and repeats.</strong> Run it by hand, daily or hourly, with several runs per check.</li>\n          </ul>\n          <p>Each run records every step: the page the agent saw, what it did and its reasoning. When a run fails, the replay opens on the step where it got stuck, and you can compare it with the last run that passed. Failing tests can alert your team by email or Slack.</p>\n          <h2>What to measure</h2>\n          <ul>\n            <li><strong>Pass rate.</strong> Because agents vary, one run proves little. Run each check several times and look at the share that passed. Ghost Agents report a journey as passing, flaky (some runs passed) or failing (none did).</li>\n            <li><strong>Where it fails.</strong> The step where failures cluster is your fix list: the size picker, the cookie banner, the account wall.</li>\n            <li><strong>Steps taken.</strong> A journey that passes in 12 steps one week and 25 the next has become harder, even if it still passes.</li>\n            <li><strong>Change over time.</strong> A journey that passed yesterday and fails today usually means something on your site changed: a theme update, a new app, a new pop-up.</li>\n          </ul>\n          <div>\n            <p><strong>Flaky is a finding, not noise.</strong> If an agent completes a journey two times in three, a real customer's assistant fails one time in three. Treat a flaky result as a problem to fix, and look at the failed runs to see what was different.</p>\n          </div>\n          <h2>Keep testing safe</h2>\n          <p>Testing with agents means letting software click buttons on your live site. Set some rules:</p>\n          <ul>\n            <li><strong>No real purchases.</strong> Stop at the payment step. Ghost Agents stop when they reach a payment form, never enter card details, and refuse buttons that place orders, pay, or delete or cancel accounts. When testing by hand, tell the assistant not to buy, and stay in control at payment.</li>\n            <li><strong>Test data only.</strong> Use a dedicated test email address, such as one on a domain you control, and dummy names and phone numbers. Never use real card or bank details, or a real customer's details.</li>\n            <li><strong>Watch your side effects.</strong> Test sign-ups create accounts, test leads land in your CRM and test bookings take slots. Use recognizable test values (a \"test+\" email or the name \"Ghost Test\") so your team can filter them out, and avoid submitting booking forms that hold real inventory.</li>\n            <li><strong>Tell the people who need to know.</strong> Let whoever runs bot protection and analytics know tests are coming. Ghost Agents identify themselves with <code>GhostAgent/1.0</code> in the user agent and sign their requests with Web Bot Auth, so they can be recognized and filtered from reports.</li>\n            <li><strong>Consider staging for risky journeys.</strong> If a journey can't be tested on the live site without real consequences, test it on a staging copy that matches production.</li>\n          </ul>\n          <h2>Turn results into fixes</h2>\n          <p>A failed run is only useful if someone acts on it. For each failure, record the journey, the step, what the agent was trying to do and a screenshot, then route it to whoever owns that part of the site: the developer for an unlabelled button, the platform admin for forced account creation, marketing for the pop-up. The replay usually tells you which.</p>\n          <p>Most fixes are the same ones covered elsewhere in these guides: <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons and forms agents can use</a>, <a href=\"https://ghostagentlab.com/articles/cookie-banners-popups-ai-agents/\">cookie banners and pop-ups</a>, <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">checkout</a> and <a href=\"https://ghostagentlab.com/articles/agent-friendly-sign-up-booking/\">sign-up and booking forms</a>.</p>\n          <h2>Where to start this week</h2>\n          <ol>\n            <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> on your home page to clear the obvious blockers first.</li>\n            <li>Pick your single most valuable journey and try it by hand with an AI assistant, three times.</li>\n            <li>Set it up as a Ghost Agent test with a daily schedule and a pass condition on the final page, so you hear about it the day it breaks.</li>\n          </ol>",
      "date_published": "2026-10-09T12:01:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Task completion"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/ai-referral-traffic/",
      "url": "https://ghostagentlab.com/articles/ai-referral-traffic/",
      "title": "Tracking visits and sales that come from AI assistants",
      "summary": "How to spot visits from ChatGPT, Perplexity and other AI assistants in your analytics, report them in GA4, and understand what attribution misses.",
      "content_html": "<p>When someone asks ChatGPT, Perplexity or Gemini for a recommendation and clicks a link in the answer, they land on your site as a visitor like any other. Those visits can be measured in the analytics you already have, and they're often worth separating out. This guide shows how to spot them, how to report on them, and what the numbers can't tell you.</p>\n          <h2>Two different things to measure</h2>\n          <p>AI shows up in your data in two ways, and it helps to keep them apart:</p>\n          <ul>\n            <li><strong>AI agents visiting your site.</strong> Crawlers and assistants that read your pages, mostly without running your analytics tag. These only show up in server or CDN logs. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure AI agent traffic</a>.</li>\n            <li><strong>People arriving from AI assistants.</strong> A person reads an AI answer, clicks a link to your site, and browses in their own browser. Your analytics tag runs normally. This is AI referral traffic, and it's what this guide covers.</li>\n          </ul>\n          <p>The first tells you whether AI can read and use your site. The second tells you whether that's turning into visits and sales.</p>\n          <h2>How AI referrals show up</h2>\n          <p>Analytics tools work out where a visit came from in two main ways: the referrer, which the browser sends to say which site the person came from, and campaign tags (UTM parameters) on the link itself.</p>\n          <h3>Referrers</h3>\n          <p>When someone clicks a link in a web-based assistant, the browser often sends that assistant's domain as the referrer. Domains you're likely to see include:</p>\n          <table>\n            <thead><tr><th>Assistant</th><th>Referrer domains you may see</th></tr></thead>\n            <tbody>\n              <tr><td>ChatGPT</td><td><code>chatgpt.com</code> (older visits may show <code>chat.openai.com</code>)</td></tr>\n              <tr><td>Perplexity</td><td><code>perplexity.ai</code>, <code>www.perplexity.ai</code></td></tr>\n              <tr><td>Google Gemini</td><td><code>gemini.google.com</code></td></tr>\n              <tr><td>Microsoft Copilot</td><td><code>copilot.microsoft.com</code></td></tr>\n              <tr><td>Claude</td><td><code>claude.ai</code></td></tr>\n            </tbody>\n          </table>\n          <p>Treat this list as a starting point rather than a complete map. Assistants change domains, launch new apps and change how they link out, so check your own referral report for new sources every month or so.</p>\n          <h3>Campaign tags</h3>\n          <p>Some assistants add their own tag to outbound links. ChatGPT, for example, has been adding <code>utm_source=chatgpt.com</code> to many links in its answers. In your analytics that shows as the source <code>chatgpt.com</code>, sometimes with no medium, or alongside the referrer. Other assistants may do the same, or nothing at all.</p>\n          <h3>Visits that arrive with no label</h3>\n          <p>Not every AI referral announces itself. Mobile and desktop apps often open links without sending a referrer, privacy settings and browsers can strip it, and a person who copies a URL from an answer and pastes it into a new tab arrives as \"direct\". Some of your direct traffic is AI referral traffic you can't see. Treat any AI referral number as a floor, not a total.</p>\n          <h2>Reporting on it in Google Analytics 4</h2>\n          <p>Analytics tools have been changing how they group this traffic. In GA4's default channel grouping, AI assistants have typically ended up under Referral, mixed in with every other website that links to you. Check what your property shows today, because default channel definitions get updated. If AI referrals aren't already separated, two approaches work:</p>\n          <ol>\n            <li><strong>A quick view.</strong> In the Traffic acquisition report, change the dimension to \"Session source\" and filter for the domains above. Good for a one-off look.</li>\n            <li><strong>A custom channel group.</strong> In Admin, under Data display, create a channel group with an \"AI assistants\" channel whose rule matches the session source against a pattern, and place it above Referral so it wins. Custom channel groups apply to your reports going back in time, so you see history too.</li>\n          </ol>\n          <p>A source pattern to start from:</p>\n          <pre><code>(^|\\.)(chatgpt\\.com|chat\\.openai\\.com|perplexity\\.ai|gemini\\.google\\.com|copilot\\.microsoft\\.com|claude\\.ai)$</code></pre>\n          <p>Use a \"matches regex\" condition on session source. Keep the pattern in one place, such as a shared doc, and add new assistants as they appear in your referral report. Other analytics tools work the same way: a rule on referrer or source that puts these visits in their own channel.</p>\n          <div>\n            <p><strong>Google's own AI features are harder.</strong> Clicks from AI Overviews and AI Mode in Google Search generally arrive as ordinary Google organic traffic, and analytics can't separate them from classic search results. Search Console reports them as part of overall Search performance. Don't expect your \"AI assistants\" channel to include them.</p>\n          </div>\n          <h2>What to measure</h2>\n          <ul>\n            <li><strong>Sessions and trend.</strong> AI referral sessions by assistant, week over week. The direction matters more than the size today.</li>\n            <li><strong>Landing pages.</strong> Which pages assistants send people to. Often it's product pages, comparison content or policy pages rather than the home page. Those pages now act as front doors, so make sure they convert.</li>\n            <li><strong>Engagement and conversion.</strong> Conversion rate, revenue and leads from the AI channel compared with search and other referrals. People arriving from an AI answer have often already been given a recommendation, so watch how their behavior compares.</li>\n            <li><strong>Assisted conversions.</strong> Someone may first hear about you from an assistant and buy a week later through a branded search. In GA4, compare attribution models or look at conversion paths to see where AI sources appear early in the journey.</li>\n          </ul>\n          <h2>The limits of attribution</h2>\n          <p>Be honest with stakeholders about what this data can and can't show:</p>\n          <ul>\n            <li><strong>It undercounts.</strong> App traffic and copied links arrive as direct. The real number is higher than what you see.</li>\n            <li><strong>It only counts clicks.</strong> Many AI answers mention a brand without a click. Someone who reads \"Northwind has the best return policy\" and later types your address directly is AI-influenced but won't appear here.</li>\n            <li><strong>It doesn't count agent purchases.</strong> When an assistant buys through an <a href=\"https://ghostagentlab.com/articles/agentic-commerce-protocols/\">agentic commerce protocol</a>, there may be no web session at all. Those orders show up in your commerce platform, so ask your platform how they're tagged.</li>\n            <li><strong>Browser agents blur the line.</strong> An agent browsing on someone's behalf may run your tag and look like a normal visitor. Some of what looks like human traffic may be agents.</li>\n          </ul>\n          <h2>Getting more AI referral traffic</h2>\n          <p>Assistants recommend and link to sites they can read and trust. The groundwork is the same as for agent readiness: let AI search and assistant agents in through <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt</a>, put key content in the HTML, describe products and prices with <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">structured data</a>, and make policy pages easy to quote. If you block <code>OAI-SearchBot</code> or <code>PerplexityBot</code>, don't expect referrals from ChatGPT search or Perplexity.</p>\n          <p>Then make sure the visits convert. The person arriving from an AI answer, or the agent acting for them, needs a page that works. <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> shows how your site looks to AI agents, and <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> test the journeys that turn a referral into a sale.</p>",
      "date_published": "2026-10-09T12:00:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Measurement"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/next-customer-not-human/",
      "url": "https://ghostagentlab.com/blog/next-customer-not-human/",
      "title": "Your next customer won't be human",
      "summary": "AI agents now read, compare and buy on people's behalf. Most websites can't see them, don't know if they succeed, and were never built for them. Here's what's changing.",
      "content_html": "<p>For twenty-five years, websites have been built for two kinds of visitor: people, and the search crawlers that send people. A third kind is arriving fast. AI agents now read, compare, book and buy on people's behalf, and most websites can't see them, don't know whether they succeed, and were never designed for them.</p>\n          <h2>A new kind of visitor</h2>\n          <p>When someone asks an AI assistant \"which of these running shoes is best for flat feet, and is it in stock in a 10?\", a person never visits your site. Software does. That software comes in three broad kinds:</p>\n          <ul>\n            <li><strong>Training crawlers</strong>, such as GPTBot and ClaudeBot, collect pages to teach AI models. They shape what AI knows about you over months.</li>\n            <li><strong>AI search and assistant fetchers</strong>, such as OAI-SearchBot, ChatGPT-User and Perplexity-User, fetch your pages in the moment to answer a question. They decide whether you're the answer today.</li>\n            <li><strong>Browser agents</strong> drive a real browser for a person: searching your catalog, filling in forms, adding to cart, and in some cases checking out.</li>\n          </ul>\n          <p>The last two are the ones that matter most to revenue, because each visit stands in for a real customer with a real intent. If the agent can't find the price, read the return policy, or get past a pop-up, the customer doesn't get a worse experience on your site. They get a recommendation for someone else's.</p>\n          <h2>Why nobody noticed</h2>\n          <p>Most websites have no idea how many agents visit, for a simple reason: analytics tools count visitors with a JavaScript tag, and most agents never run it. Crawlers and assistant fetchers request the HTML and leave. Your dashboards show a quiet day while your server logs show a busy one.</p>\n          <p>The agents that do run JavaScript often look like ordinary browsers, so they blend into the human numbers. Either way, the question \"did the agent get what it came for?\" goes unanswered.</p>\n          <h2>What goes wrong</h2>\n          <p>When we test sites with real agents, the failures fall into the same few patterns:</p>\n          <ol>\n            <li><strong>They're turned away at the door.</strong> A robots.txt rule written years ago, or bot protection that challenges anything that isn't a person, blocks the assistants that were trying to send customers.</li>\n            <li><strong>They can't read the page.</strong> Prices and product details only appear after JavaScript runs, and there's no structured data to fall back on.</li>\n            <li><strong>They can't find their way.</strong> Unlabelled icon buttons, menus that only open on hover, and cookie banners that cover the page stop an agent the way a locked door would.</li>\n            <li><strong>They can't finish the job.</strong> The agent finds the product but can't choose a size, or reaches the cart but can't find the checkout button.</li>\n          </ol>\n          <p>None of these show up in a normal QA pass, because a person gets through every one of them without noticing.</p>\n          <h2>The gap in today's tools</h2>\n          <p>Two kinds of tools touch this problem, and neither solves it.</p>\n          <p><strong>Bot management</strong> decides who to block. It's essential, and it's built to keep bad traffic out, which means good agents are often collateral damage. <strong>AI visibility</strong> tools track what chatbots say about your brand. That's useful too, but it measures the answer, not whether an agent could use your site to get there.</p>\n          <p>What's missing is the thing we've always done for human visitors: test the experience. Can an agent actually complete the journeys that make you money, and does it still work after this week's release?</p>\n          <h2>Agent experience</h2>\n          <p>We think agent experience will become a discipline of its own, next to user experience and SEO. It has three parts, and they're how we've built Ghost Agent Labs:</p>\n          <ul>\n            <li><strong>Observe.</strong> See every crawler, assistant and browser agent that visits, from server and edge data rather than a JavaScript tag, and check which ones are real.</li>\n            <li><strong>Test.</strong> Send real AI agents through checkout, sign-up and search on a schedule, and get alerted when they fail where a person would succeed.</li>\n            <li><strong>Improve.</strong> Score agent readiness, fix the highest-impact problems first, and re-test to prove the fix worked.</li>\n          </ul>\n          <p>It's the same idea behind <a href=\"https://ghostinspector.com/\">Ghost Inspector</a>, which has tested websites for human visitors for years, applied to the visitors that aren't human.</p>\n          <h2>What to do this week</h2>\n          <div>\n            <ol>\n              <li>Read your robots.txt and check you aren't blocking AI assistants or AI search by accident. Our guide to <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a> walks through it.</li>\n              <li>Ask your CDN or bot protection vendor how it treats verified AI agents, and whether it can tell real ones from fakes. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</li>\n              <li>Add an <a href=\"https://ghostagentlab.com/articles/llms-txt/\">llms.txt file</a> that points agents to your most important pages.</li>\n              <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>. In under a minute, it shows how AI agents are let in, what they can read and what gets in their way.</li>\n            </ol>\n          </div>\n          <p>Your next million visitors won't be human. The sites that are ready for them will win the customers they bring.</p>",
      "date_published": "2026-10-08T12:11:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Perspective"
      ]
    },
    {
      "id": "https://ghostagentlab.com/blog/signed-requests/",
      "url": "https://ghostagentlab.com/blog/signed-requests/",
      "title": "Why we sign every request our agents send",
      "summary": "Anyone can copy a user agent. Here's how Ghost Agent Labs signs its scanner and Ghost Agent traffic with Web Bot Auth, and why every well-behaved agent should.",
      "content_html": "<p>A user agent string is a name tag anyone can print. If AI agents are going to shop, book and sign up on people's behalf, websites need a better way to tell who's knocking. That's why every request our scanner and Ghost Agents send under their own name is cryptographically signed.</p>\n          <h2>The problem with user agents</h2>\n          <p>Every request a browser or bot makes carries a <code>User-Agent</code> header. Ours look like this:</p>\n          <pre><code>AgentScore/1.0 (+https://ghostagentlab.com/agentscore/bot)\nGhostAgent/1.0 (+https://ghostagentlab.com/ghost-agent)</code></pre>\n          <p>That's useful, and it's a courtesy we'll keep. But it proves nothing. A scraper can send <code>GPTBot</code>, <code>ClaudeBot</code> or <code>AgentScore/1.0</code> just as easily as the real thing, and plenty do. So site owners face a bad choice: trust the label and let impostors in, or block the label and turn away the agents they actually want.</p>\n          <p>The traditional fix is to check where the request came from: published IP ranges, or a reverse DNS lookup that resolves back to the operator's domain. That works for big crawlers with fixed infrastructure. It works badly for browser agents running in the cloud, where IP addresses change all the time, and it puts the burden on every website to keep lists up to date.</p>\n          <h2>What Web Bot Auth does instead</h2>\n          <p><a href=\"https://blog.cloudflare.com/web-bot-auth/\">Web Bot Auth</a> is a proposal, now being worked on at the IETF and supported by Cloudflare's verified bots program, that lets an agent prove who it is the same way websites prove who they are: with a public key.</p>\n          <ol>\n            <li>The agent operator publishes its public keys in a directory at a well-known address on its own domain. Ours is <code>https://app.ghostagentlab.com/.well-known/http-message-signatures-directory</code>.</li>\n            <li>Every request carries three extra headers: <code>Signature-Agent</code> (where to find the keys), <code>Signature-Input</code> (what was signed, when, and with which key) and <code>Signature</code> (the signature itself), following HTTP Message Signatures, <a href=\"https://www.rfc-editor.org/rfc/rfc9421\">RFC 9421</a>.</li>\n            <li>The website or its CDN fetches the key once, checks the signature, and knows for certain the request came from the operator, whatever IP address it arrived from.</li>\n          </ol>\n          <p>Anyone can copy a user agent. Only the holder of the private key can produce the signature.</p>\n          <h2>How we do it</h2>\n          <p>We sign with an Ed25519 key. Each signature covers the site's host name and our <code>Signature-Agent</code> header, carries a fresh random nonce, and expires after 60 seconds, so a captured signature can't be replayed against another site or reused later. The key ID is the key's standard JWK thumbprint, and the key directory itself is signed too, so nobody can swap in their own keys.</p>\n          <table>\n            <thead><tr><th>Requests</th><th>Signed?</th></tr></thead>\n            <tbody>\n              <tr><td>AgentScore requests under its own user agent: robots.txt, sitemaps, well-known files, page discovery</td><td>Yes</td></tr>\n              <tr><td>Every Ghost Agent request, including images and scripts from other hosts</td><td>Yes</td></tr>\n              <tr><td>AgentScore requests that deliberately imitate a browser or another agent, such as ChatGPT-User, to see how your site treats them</td><td>No, so your site treats them exactly as it would that visitor</td></tr>\n            </tbody>\n          </table>\n          <p>That last row matters. Part of what AgentScore measures is whether your bot protection treats AI agents differently from people. If we signed those requests, a CDN that trusts us would wave them through and the test would tell you nothing.</p>\n          <h2>Why this matters beyond us</h2>\n          <p>Agents are becoming customers' representatives. When someone asks an assistant to reorder printer ink or book a table, the agent that arrives at your site is acting for a real person with a real wallet. The sites that win are the ones that can say yes to those agents confidently, without opening the door to every scraper wearing a borrowed name.</p>\n          <p>Signed requests make that possible. A site can allow verified agents through a challenge page, give them a lighter rate limit, or simply count them accurately in analytics. None of that works if the identity is a string anyone can type.</p>\n          <div>\n            <p><strong>If you run an agent:</strong> sign your requests. Publish a key directory, add the three headers, and register with the CDNs your users' sites sit behind. It's a small amount of work, and it's the difference between being trusted and being blocked.</p>\n            <p><strong>If you run a website:</strong> check whether your CDN or bot protection can verify Web Bot Auth signatures, and prefer it over IP allowlists for the agents you want. Then run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> to see how your site treats AI agents today.</p>\n          </div>\n          <h2>Checking it's really us</h2>\n          <p>The details of what our agents do, what they won't do, and how to opt out are on the <a href=\"https://ghostagentlab.com/agentscore/bot/\">AgentScore bot</a> and <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> pages. We don't use fixed IP addresses yet, so if you want to allow us, please do it by signature or user agent rather than by IP.</p>",
      "date_published": "2026-10-08T12:10:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Blog",
        "Engineering"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/agent-readiness/",
      "url": "https://ghostagentlab.com/articles/agent-readiness/",
      "title": "What is agent readiness? A complete guide",
      "summary": "How well AI agents can get into your site, understand it, find their way around and complete tasks: what each part means, what good looks like, and where to start.",
      "content_html": "<p>Agent readiness is how well AI agents can get into your website, understand it, find their way around it, and complete the tasks people send them to do. This guide explains each part, what good looks like, and where to start.</p>\n          <h2>Why it matters now</h2>\n          <p>AI assistants increasingly visit websites on a person's behalf: to answer a question, compare products, check a policy, or buy something. Each of those visits stands in for a customer. If the agent can't use your site, it doesn't complain or call support. It moves on to a site it can use, and recommends that one instead.</p>\n          <p>Agent readiness is not the same as SEO, though they overlap. SEO is about being found and ranked. Agent readiness is about what happens next: whether an agent that has found you can actually do something useful with your site.</p>\n          <p>Not every agent wants the same thing. A crawler collecting pages for AI training is very different from an assistant checking your returns policy for a customer, or a browser agent trying to check out. See <a href=\"https://ghostagentlab.com/articles/kinds-of-ai-agents/\">the AI agents that visit your website</a>, and keep the <a href=\"https://ghostagentlab.com/articles/agent-readiness-glossary/\">agent readiness glossary</a> handy for the terms in this guide.</p>\n          <h2>The four parts of agent readiness</h2>\n          <p><a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> measures agent readiness out of 100, across four categories. The weights reflect what stops agents most often in practice.</p>\n          <table>\n            <thead><tr><th>Category</th><th>The question</th><th>Points</th></tr></thead>\n            <tbody>\n              <tr><td>Access</td><td>Can agents get in?</td><td>25</td></tr>\n              <tr><td>Readability</td><td>Can agents understand your pages?</td><td>25</td></tr>\n              <tr><td>Navigability</td><td>Can agents find their way around?</td><td>20</td></tr>\n              <tr><td>Task completion</td><td>Can an agent actually get the job done?</td><td>30</td></tr>\n            </tbody>\n          </table>\n          <h3>1. Access: can agents get in?</h3>\n          <p>Before anything else, the agent has to be allowed through the door. Most access failures are accidental: a robots.txt rule copied from a template, or bot protection that challenges every visitor that isn't a person.</p>\n          <ul>\n            <li><strong>robots.txt</strong> allows AI assistants and AI search agents. You can still block training crawlers separately. See <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a>.</li>\n            <li><strong>Bot protection</strong> doesn't block AI agents that a person would get through, and can tell real agents from impostors. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a> and <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</li>\n            <li><strong>No CAPTCHA or challenge page</strong> on arrival for ordinary pages.</li>\n            <li><strong>An llms.txt file</strong> points agents to your most important pages. See <a href=\"https://ghostagentlab.com/articles/llms-txt/\">how to write an llms.txt file</a>.</li>\n          </ul>\n          <h3>2. Readability: can agents understand your pages?</h3>\n          <p>Many agents read the HTML your server sends and never run JavaScript. If your prices, product details or policies only appear after scripts load, those agents see an empty page.</p>\n          <ul>\n            <li><strong>Key content is in the HTML</strong>, through server-side rendering or static generation, not only added by JavaScript. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see JavaScript-only content</a>.</li>\n            <li><strong>Structured data</strong> (schema.org) describes products, prices, availability and your business. See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data AI shopping agents can read</a>.</li>\n            <li><strong>Clear titles, headings and image descriptions</strong> on every important page.</li>\n            <li><strong>A valid sitemap</strong> that lists the pages you want found.</li>\n          </ul>\n          <h3>3. Navigability: can agents find their way around?</h3>\n          <p>Browser agents find their way around a page much as a screen reader does: by the names and roles of buttons, links and form fields. Much of agent readiness here is simply good accessibility. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>.</p>\n          <ul>\n            <li><strong>Every button, link and form field has a name</strong> an agent can read. An icon with no label is invisible.</li>\n            <li><strong>Cookie banners and pop-ups</strong> can be dismissed and don't cover the content.</li>\n            <li><strong>Menus open on click or focus</strong>, not only on hover.</li>\n            <li><strong>Page landmarks</strong> such as header, navigation, main and footer let agents find what they need quickly.</li>\n          </ul>\n          <h3>4. Task completion: can an agent get the job done?</h3>\n          <p>This is the part that matters most, and the only way to measure it is to try. A site can pass every technical check and still lose agents at the size selector or the checkout button.</p>\n          <ul>\n            <li><strong>Online stores:</strong> an agent can search, find a product, choose options and add it to the cart. See <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">how to make checkout work for AI shopping agents</a>.</li>\n            <li><strong>Software companies:</strong> an agent can find pricing and reach the sign-up form.</li>\n            <li><strong>Service businesses:</strong> an agent can find prices, availability and a way to book.</li>\n          </ul>\n          <p>The free AgentScore scan doesn't test this yet: it marks Task completion “not tested” and works out your score from the other three categories. <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> test it today, running the journeys you choose on a schedule, with a step-by-step replay of each run, and alerting you when one breaks.</p>\n          <h2>How to measure your agent readiness</h2>\n          <ol>\n            <li><strong>Run AgentScore.</strong> It's free, takes under a minute, and checks Access, Readability and Navigability. Task completion isn't part of the free scan yet.</li>\n            <li><strong>Look at your agent traffic.</strong> Your analytics probably can't see agents, because most never run the JavaScript tag. Server or CDN logs can, and Ghost Agent Labs reads them for you. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure agent traffic</a>.</li>\n            <li><strong>Test your money journeys.</strong> Checkout, sign-up and booking are the journeys where an agent failure costs you a customer.</li>\n          </ol>\n          <h2>A 30-minute starter checklist</h2>\n          <div>\n            <ol>\n              <li>Open <code>/robots.txt</code> and make sure it doesn't block AI assistants or AI search agents.</li>\n              <li>Ask whoever runs your CDN or bot protection how it treats verified AI agents.</li>\n              <li>View the source of a product or pricing page (not the inspector) and check the price is in the HTML.</li>\n              <li>Publish an llms.txt file with links to your key pages.</li>\n              <li>Tab through your header with the keyboard. If you can't reach the menu or the cart, neither can many agents.</li>\n              <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> and fix the top three items it lists.</li>\n            </ol>\n          </div>\n          <h2>Common questions</h2>\n          <h3>Won't letting agents in mean letting scrapers in?</h3>\n          <p>No. Agent readiness is about letting in the agents you want, and the main ones can be verified. Keep blocking unknown and abusive bots; just make sure your rules don't catch AI assistants that are bringing you customers.</p>\n          <h3>Do I have to let AI companies train on my content?</h3>\n          <p>No. Training crawlers and the agents that fetch pages for people use different names, so you can block one and allow the other. Our <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt guide</a> shows how.</p>\n          <h3>Is this just accessibility?</h3>\n          <p>There's a big overlap, especially in navigation, and work on one helps the other. But agent readiness also covers access rules, machine-readable data, and whether an agent can finish a real task.</p>",
      "date_published": "2026-10-08T12:09:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Start here"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/robots-txt-ai-agents/",
      "url": "https://ghostagentlab.com/articles/robots-txt-ai-agents/",
      "title": "robots.txt for AI agents: who to allow and who to block",
      "summary": "The three kinds of AI agent, the robots.txt names each one uses, a sensible default, and the most-specific-group rule that catches most people out.",
      "content_html": "<p>Your robots.txt file may be blocking the AI assistants that send you customers, while doing nothing to stop the bots you actually wanted to keep out. This guide explains the three kinds of AI agent, which robots.txt names each one uses, and how to write rules that let the right ones in.</p>\n          <h2>Three kinds of AI agent</h2>\n          <p>\"AI bots\" isn't one thing. The major AI companies each run several agents with different jobs and different names, and the right policy for each is different.</p>\n          <table>\n            <thead><tr><th>Kind</th><th>What it does</th><th>Examples (robots.txt name)</th></tr></thead>\n            <tbody>\n              <tr><td>Training crawlers</td><td>Collect pages to train AI models</td><td><code>GPTBot</code>, <code>ClaudeBot</code>, <code>CCBot</code>, <code>Bytespider</code>, <code>meta-externalagent</code></td></tr>\n              <tr><td>AI search crawlers</td><td>Index pages so AI search can cite and link to them</td><td><code>OAI-SearchBot</code>, <code>Claude-SearchBot</code>, <code>PerplexityBot</code>, <code>Amazonbot</code></td></tr>\n              <tr><td>Assistant fetchers</td><td>Fetch a page in real time because a person asked about it</td><td><code>ChatGPT-User</code>, <code>Claude-User</code>, <code>Perplexity-User</code>, <code>meta-externalfetcher</code>, <code>DuckAssistBot</code>, <code>MistralAI-User</code></td></tr>\n            </tbody>\n          </table>\n          <p>Some names aren't crawlers at all. <code>Google-Extended</code> and <code>Applebot-Extended</code> are control tokens: Google and Apple crawl with Googlebot and Applebot as usual, and these names let you say whether that content may be used for their AI models. Blocking them doesn't affect normal search.</p>\n          <p>Lists like this change as companies launch new agents, so check each operator's documentation for the current names.</p>\n          <h2>A sensible default</h2>\n          <p>For most businesses, the agents that matter most are AI search and assistant fetchers, because they answer people's questions about you and link back to you. Training is a separate decision about your content, and it's yours to make.</p>\n          <p>This robots.txt allows everything, which is also what you get with no robots.txt at all:</p>\n          <pre><code>User-agent: *\nAllow: /\nSitemap: https://www.example.com/sitemap.xml</code></pre>\n          <p>If you'd rather not have your content used to train AI models, but still want to appear in AI search and assistant answers, block only the training names:</p>\n          <pre><code># Allow AI search and assistants; opt out of AI training\nUser-agent: GPTBot\nUser-agent: ClaudeBot\nUser-agent: CCBot\nUser-agent: Google-Extended\nUser-agent: Applebot-Extended\nUser-agent: meta-externalagent\nDisallow: /\nUser-agent: *\nDisallow: /cart\nDisallow: /account\nAllow: /\nSitemap: https://www.example.com/sitemap.xml</code></pre>\n          <p>Listing several <code>User-agent</code> lines above one set of rules applies those rules to all of them.</p>\n          <h2>The rule that catches most people out</h2>\n          <p>A crawler follows only the most specific group that matches its name, and ignores the rest. If there's a group for <code>GPTBot</code>, GPTBot ignores everything under <code>User-agent: *</code>.</p>\n          <p>So in the example above, the training crawlers don't inherit the <code>/cart</code> and <code>/account</code> rules; they don't need them, because they're blocked from everything. But the reverse mistake is common:</p>\n          <pre><code>User-agent: *\nDisallow: /checkout\nDisallow: /admin\nUser-agent: OAI-SearchBot\nAllow: /</code></pre>\n          <p>Someone added that second group to \"make sure\" OpenAI's search crawler gets in. Instead, they've told it that <code>/checkout</code> and <code>/admin</code> are open, because it no longer reads the <code>*</code> group. If you add a group for a specific agent, copy into it every rule you want it to follow.</p>\n          <h2>Other common mistakes</h2>\n          <ul>\n            <li><strong>A leftover <code>Disallow: /</code>.</strong> Staging sites are often launched with a robots.txt that blocks everything. Check yours today.</li>\n            <li><strong>Blocking every AI name from a copied list.</strong> Long \"block all AI\" lists usually include the assistant fetchers and AI search crawlers too, which removes you from AI answers as well as training.</li>\n            <li><strong>Blocking the pages agents need.</strong> Disallowing <code>/products</code> or <code>/pricing</code> for all bots, or for <code>*</code>, to save crawl budget also hides them from agents.</li>\n            <li><strong>Relying on robots.txt to block bad bots.</strong> robots.txt is a request, not a lock. Well-behaved crawlers follow it; scrapers ignore it. Abusive traffic has to be stopped at your CDN or firewall.</li>\n            <li><strong>Forgetting your CDN.</strong> A firewall rule or bot setting that challenges \"AI crawlers\" overrides a friendly robots.txt. The agent never gets far enough to read it.</li>\n          </ul>\n          <h2>A note on assistant fetchers</h2>\n          <p>When a person asks an assistant to read a specific page, some operators treat that fetch like a person clicking a link rather than a crawl, and say robots.txt rules may not apply to it. If you need to stop those visits completely, that has to happen at your CDN or server, not in robots.txt. For most businesses, though, these are the visits you want most: each one is a customer asking about you.</p>\n          <h2>How to check your robots.txt</h2>\n          <ol>\n            <li>Open <code>https://yourdomain.com/robots.txt</code> and read it top to bottom. Note every group and which names it applies to.</li>\n            <li>For each AI search crawler and assistant fetcher in the table above, work out which group it follows, remembering the most-specific-group rule, and whether that group blocks your important pages.</li>\n            <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>. Its robots.txt check reads your rules as 21 AI agents would and tells you which assistants and AI search agents are blocked from your home page. Blocked training crawlers are reported but don't cost you points, because that's a legitimate choice.</li>\n          </ol>\n          <p>robots.txt is only the first door. Next, make sure your bot protection lets the real agents through, and can tell them apart from impostors using their names. See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>.</p>",
      "date_published": "2026-10-08T12:08:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/verify-ai-crawlers/",
      "url": "https://ghostagentlab.com/articles/verify-ai-crawlers/",
      "title": "How to tell if an AI crawler is real",
      "summary": "Anyone can claim to be GPTBot or Googlebot. How to verify AI agents with published IP ranges, reverse DNS and signed requests, and what to do when there's no proof.",
      "content_html": "<p>Anyone can send a request that says it's GPTBot or Googlebot. To let real AI agents in while keeping impostors out, you need to check where a request actually came from. There are three ways to do it, and the right one depends on the operator.</p>\n          <h2>Why user agents can't be trusted</h2>\n          <p>Every request names its sender in the <code>User-Agent</code> header, and that name is a promise, not a proof. Scrapers routinely borrow the names of well-known crawlers, because many sites let those crawlers through without a challenge. If your bot rules trust the name alone, you're letting in anyone who copies it. If they distrust it, you're blocking real agents along with the fakes.</p>\n          <p>Verification resolves that. It answers one question: did this request really come from the company it names?</p>\n          <h2>Method 1: published IP ranges</h2>\n          <p>Several operators publish the IP addresses their crawlers use, as machine-readable JSON files. If a request claims to be one of their agents and comes from an address in the list, it's real.</p>\n          <table>\n            <thead><tr><th>Operator</th><th>Agents covered</th></tr></thead>\n            <tbody>\n              <tr><td>Google</td><td>Googlebot, special-case crawlers and user-triggered fetchers (separate lists)</td></tr>\n              <tr><td>Microsoft</td><td>Bingbot</td></tr>\n              <tr><td>Apple</td><td>Applebot</td></tr>\n              <tr><td>OpenAI</td><td>GPTBot, OAI-SearchBot and ChatGPT-User (one list each)</td></tr>\n              <tr><td>Perplexity</td><td>PerplexityBot and Perplexity-User (one list each)</td></tr>\n            </tbody>\n          </table>\n          <p>Each operator links its list from its crawler documentation. A few practical points:</p>\n          <ul>\n            <li><strong>Refresh the lists often.</strong> Ranges change. Fetch them at least daily, and don't treat a request as fake because it's missing from a list that's days old.</li>\n            <li><strong>Match the list to the agent.</strong> An OpenAI address in the GPTBot list doesn't verify a request claiming to be ChatGPT-User.</li>\n            <li><strong>Handle IPv6.</strong> Several lists include IPv6 ranges.</li>\n          </ul>\n          <h2>Method 2: forward-confirmed reverse DNS</h2>\n          <p>Some operators instead promise that their crawlers' IP addresses resolve to a host name on their own domain. The check has two steps, because reverse DNS on its own can be faked by whoever controls the IP address:</p>\n          <ol>\n            <li><strong>Reverse lookup:</strong> look up the host name for the request's IP address, and check it ends in the operator's domain.</li>\n            <li><strong>Forward lookup:</strong> look up that host name's IP addresses, and check the original IP is among them.</li>\n          </ol>\n          <figure>\n            <ol>\n              <li>\n                <p>A request arrives</p>\n                <p>User agent says Googlebot, from IP address <code>66.249.66.1</code></p>\n              </li>\n              <li>\n                <p>1. Reverse lookup</p>\n                <p><code>66.249.66.1</code> &rarr; <code>crawl-66-249-66-1.googlebot.com</code></p>\n                <p>Ends in googlebot.com</p>\n              </li>\n              <li>\n                <p>2. Forward lookup</p>\n                <p><code>crawl-66-249-66-1.googlebot.com</code> &rarr; <code>66.249.66.1</code></p>\n                <p>Points back to the same address</p>\n              </li>\n              <li>\n                <p>Verified</p>\n                <p>The request is from Google. If either step fails, it isn&rsquo;t, whatever its user agent says.</p>\n              </li>\n            </ol>\n            <figcaption>Forward-confirmed reverse DNS, using Google&rsquo;s own example. The second step matters because whoever controls an IP address can make its reverse lookup say anything.</figcaption>\n          </figure>\n          <p>Google's own example, from a terminal:</p>\n          <pre><code>$ host 66.249.66.1\n1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.\n$ host crawl-66-249-66-1.googlebot.com\ncrawl-66-249-66-1.googlebot.com has address 66.249.66.1</code></pre>\n          <p>The name ends in <code>googlebot.com</code> and points back to the same address, so the request is from Google. Domains that operators document for this include:</p>\n          <table>\n            <thead><tr><th>Agent</th><th>Host name ends in</th></tr></thead>\n            <tbody>\n              <tr><td>Googlebot and other Google crawlers</td><td><code>googlebot.com</code>, <code>google.com</code> or <code>googleusercontent.com</code></td></tr>\n              <tr><td>Bingbot</td><td><code>search.msn.com</code></td></tr>\n              <tr><td>Applebot</td><td><code>applebot.apple.com</code></td></tr>\n              <tr><td>Amazonbot</td><td><code>crawl.amazonbot.amazon</code></td></tr>\n            </tbody>\n          </table>\n          <p>DNS lookups are slow compared with a page request, so cache the result for each IP address rather than checking on every hit.</p>\n          <h2>Method 3: signed requests (Web Bot Auth)</h2>\n          <p>The newest method doesn't depend on IP addresses at all. With <a href=\"https://blog.cloudflare.com/web-bot-auth/\">Web Bot Auth</a>, the agent signs each request with a private key and publishes the matching public key on its own domain. Your CDN or server checks the signature, and a valid one proves the sender, whatever network it came from.</p>\n          <p>That makes it the best fit for browser agents running in the cloud, whose addresses change constantly. It's still early, so check whether your CDN supports it. We explain how it works, and why we sign all of our own agents' requests, in <a href=\"https://ghostagentlab.com/blog/signed-requests/\">Why we sign every request our agents send</a>.</p>\n          <h2>When there's no proof at all</h2>\n          <p>Not every operator publishes IP ranges, a reverse DNS domain, or signing keys for every agent. For those, you can't prove a request is real or fake; you can only weigh the evidence, such as how it behaves and how fast it requests pages.</p>\n          <p>Treat \"can't tell\" as its own answer, not as \"fake\". Blocking everything you can't verify will block legitimate agents from operators that simply haven't published proof yet.</p>\n          <h2>Putting it together</h2>\n          <ol>\n            <li><strong>Identify:</strong> match the user agent to a known agent and its operator.</li>\n            <li><strong>Verify:</strong> use the strongest proof that operator offers: a signature, then published IP ranges, then reverse DNS.</li>\n            <li><strong>Decide:</strong> let verified agents you want through; challenge or block requests that claim a name but fail verification; and handle \"can't tell\" with ordinary rate limits rather than a hard block.</li>\n          </ol>\n          <div>\n            <p><strong>Your CDN may already do this.</strong> Many CDNs and bot-management products maintain lists of verified bots. Check that the AI assistants and AI search agents you want are in the allowed categories, and that \"AI crawler\" blocking rules aren't catching them by name.</p>\n            <p><strong>Ghost Agent Labs does it for you.</strong> It identifies every agent in your server or edge traffic, checks it against published IP ranges and reverse DNS, and flags impostors, so you can see which agents are real before you decide what to allow. Missing or stale data is reported as \"can't tell\", never as \"spoofed\". <a href=\"https://app.ghostagentlab.com/signup\">Start free</a>.</p>\n          </div>",
      "date_published": "2026-10-08T12:07:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/bot-protection-ai-agents/",
      "url": "https://ghostagentlab.com/articles/bot-protection-ai-agents/",
      "title": "Bot protection and CAPTCHAs: stop blocking the agents you want",
      "summary": "Why bot protection catches AI assistants, how to tell if it's happening, and how to tune your CDN, WAF and CAPTCHAs so good agents get through and bad bots don't.",
      "content_html": "<p>Bot protection exists to keep out scrapers, credential stuffers and fraud. But the same rules often stop the AI assistants trying to answer a customer's question about your products. This guide shows how to tell whether that's happening, and how to tune your defenses so the agents you want get through and the bots you don't still get stopped.</p>\n          <h2>Why good agents get caught</h2>\n          <p>Bot protection, whether it's a CDN feature, a web application firewall (WAF) or a dedicated product, looks for signs that a visitor isn't a person. AI agents show many of those signs, for entirely legitimate reasons:</p>\n          <ul>\n            <li><strong>They say they're bots.</strong> Honest agents announce themselves in their user agent, and a rule that blocks \"bots\" or \"AI crawlers\" by name catches them first.</li>\n            <li><strong>They don't run JavaScript.</strong> Many challenges work by running a script in the visitor's browser. Assistant fetchers and AI crawlers usually read only the HTML, so they never pass.</li>\n            <li><strong>They come from data centers.</strong> Agents run in the cloud, and cloud IP addresses score as higher risk than home broadband.</li>\n            <li><strong>Browser agents look automated.</strong> Even agents that drive a real browser move and type differently from people, and fingerprinting tools notice.</li>\n            <li><strong>One-click \"block AI\" settings.</strong> Several providers offer a single switch to block AI bots. Depending on how it's set up, it can block assistants and AI search along with training crawlers.</li>\n          </ul>\n          <h2>Signs it's happening to you</h2>\n          <ul>\n            <li>AI assistants say they \"can't access\" your site, or describe it from out-of-date information.</li>\n            <li>Your CDN or firewall logs show 403 (forbidden), 429 (too many requests) or challenge responses for agents like ChatGPT-User, Claude-User or Perplexity-User.</li>\n            <li>You appear in AI search results much less often than competitors with weaker SEO.</li>\n          </ul>\n          <p>The quickest test is <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>. It requests your home page as a normal browser and as five AI agents, including ChatGPT-User, Claude-User, Perplexity-User and OAI-SearchBot, and compares what comes back. It fails the check if an agent is blocked or challenged while the browser gets through, and warns if an agent gets less than half the content the browser does. Separate checks look for CAPTCHAs on your home page, and test whether agents can open your product, pricing and cart pages.</p>\n          <h2>How to fix it</h2>\n          <h3>1. Allow verified agents, not names</h3>\n          <p>The safe way to let agents in is by verified identity. Most bot management products keep a list of verified bots, checked against the operator's published IP ranges, reverse DNS or signatures, and let you allow categories of them. Allow the categories for AI assistants and AI search, and keep blocking requests that only claim those names. Our guide to <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">telling real AI crawlers from fakes</a> explains how verification works.</p>\n          <div>\n            <p><strong>Never allow by user agent alone.</strong> A rule like \"if the user agent contains GPTBot, skip all checks\" is the first thing scrapers exploit.</p>\n          </div>\n          <h3>2. Separate training from assistants</h3>\n          <p>If you've turned on a \"block AI bots\" setting, check exactly which agents it covers. If your goal is to opt out of AI training, block only training crawlers, and leave assistant fetchers and AI search crawlers allowed. See <a href=\"https://ghostagentlab.com/articles/robots-txt-ai-agents/\">robots.txt for AI agents</a> for which names are which.</p>\n          <h3>3. Protect actions, not pages</h3>\n          <p>Abuse mostly targets actions: logging in, creating accounts, applying discount codes, checking gift card balances, submitting payment. It rarely needs protecting against on pages that only display information. Concentrate strict rules and challenges on:</p>\n          <ul>\n            <li>Login, sign-up and password reset forms</li>\n            <li>Checkout submission and payment</li>\n            <li>Gift card, coupon and stock-check endpoints that attackers hammer</li>\n            <li>Search and APIs, with rate limits rather than outright blocks</li>\n          </ul>\n          <p>Keep product, category, pricing, policy and help pages as open as you safely can. Those are the pages agents need to answer questions about you.</p>\n          <h3>4. Use rate limits instead of blocks</h3>\n          <p>A real assistant fetcher makes a handful of requests because a person asked about you. A scraper makes thousands. A sensible rate limit per IP address or verified agent stops the scraper without punishing the assistant, where a blanket block stops both.</p>\n          <h3>5. Replace puzzle CAPTCHAs on key pages</h3>\n          <p>AI agents can't solve CAPTCHAs, and aren't meant to. If you need a challenge, use invisible or risk-based challenges that only escalate when traffic looks abusive, and keep them off the pages agents need to read. Never put a CAPTCHA in front of ordinary content on arrival.</p>\n          <h3>6. Don't serve agents a different page</h3>\n          <p>Some setups don't block agents outright, but quietly serve them a stripped-down page, a placeholder, or an \"enable JavaScript\" message. To the agent this is as bad as a block, and it's harder to notice. Agents should get the same content a browser does.</p>\n          <h2>A quick checklist</h2>\n          <div>\n            <ol>\n              <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a> and look at the bot protection, CAPTCHA and key pages checks.</li>\n              <li>In your CDN or bot management settings, find the verified bots or AI categories, and confirm AI assistants and AI search are allowed.</li>\n              <li>Check that any \"block AI\" setting covers only the agents you mean to block.</li>\n              <li>Search your firewall rules for user agent matches and make sure none allow access by name alone.</li>\n              <li>Move strict challenges to login, sign-up and checkout submission, and use rate limits elsewhere.</li>\n              <li>Re-run AgentScore to confirm the fix.</li>\n            </ol>\n          </div>\n          <p>To keep an eye on this over time, watch how often agents are blocked in your traffic. See <a href=\"https://ghostagentlab.com/articles/measure-ai-agent-traffic/\">how to measure AI agent traffic</a>.</p>",
      "date_published": "2026-10-08T12:06:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/",
      "url": "https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/",
      "title": "Product data that AI shopping agents can read",
      "summary": "How to describe products, prices, stock, shipping and returns with schema.org structured data, with a complete JSON-LD example and the rules that keep it trustworthy.",
      "content_html": "<p>When an AI agent compares products for someone, it needs facts it can trust: the exact name, the price, the currency, whether it's in stock, how shipping and returns work. Structured data hands it those facts in a standard format, instead of making it guess from your page layout.</p>\n          <h2>What structured data is</h2>\n          <p>Structured data is a small block of machine-readable information in your page's HTML that describes what the page is about, using the shared vocabulary at <a href=\"https://schema.org/\">schema.org</a>. The most common format is JSON-LD: a <code>&lt;script type=\"application/ld+json\"&gt;</code> tag that people never see, but search engines and agents read directly.</p>\n          <p>Search engines have used it for years to show prices and star ratings in results. For AI agents it matters even more. A person can glance at a page and see that \"$24\" next to the \"Add to cart\" button is the price. An agent reading the raw page has to infer that, and on a busy page with sale prices, bundle prices and \"from\" prices, it can get it wrong. Structured data removes the guesswork.</p>\n          <h2>A complete product example</h2>\n          <p>Here's JSON-LD for a single product page, with the details agents use most:</p>\n          <pre><code>&lt;script type=\"application/ld+json\"&gt;\n{\n  \"@context\": \"https://schema.org\",\n  \"@type\": \"Product\",\n  \"name\": \"Ethiopia Yirgacheffe Whole Bean Coffee, 340 g\",\n  \"description\": \"Light roast single-origin coffee with notes of jasmine and lemon.\",\n  \"sku\": \"NW-ETH-340\",\n  \"gtin13\": \"0123456789012\",\n  \"brand\": { \"@type\": \"Brand\", \"name\": \"Northwind Coffee\" },\n  \"image\": \"https://northwind.example/images/ethiopia-340.jpg\",\n  \"url\": \"https://northwind.example/products/ethiopia-yirgacheffe\",\n  \"aggregateRating\": {\n    \"@type\": \"AggregateRating\",\n    \"ratingValue\": \"4.7\",\n    \"reviewCount\": \"212\"\n  },\n  \"offers\": {\n    \"@type\": \"Offer\",\n    \"price\": \"18.00\",\n    \"priceCurrency\": \"USD\",\n    \"availability\": \"https://schema.org/InStock\",\n    \"itemCondition\": \"https://schema.org/NewCondition\",\n    \"url\": \"https://northwind.example/products/ethiopia-yirgacheffe\",\n    \"hasMerchantReturnPolicy\": {\n      \"@type\": \"MerchantReturnPolicy\",\n      \"applicableCountry\": \"US\",\n      \"returnPolicyCategory\": \"https://schema.org/MerchantReturnFiniteReturnWindow\",\n      \"merchantReturnDays\": 30\n    }\n  }\n}\n&lt;/script&gt;</code></pre>\n          <h2>The fields that matter most</h2>\n          <table>\n            <thead><tr><th>Field</th><th>Why agents need it</th></tr></thead>\n            <tbody>\n              <tr><td><code>name</code></td><td>The exact product, including size or pack where it matters</td></tr>\n              <tr><td><code>offers.price</code> and <code>priceCurrency</code></td><td>The price the customer will actually pay, as a plain number, with no currency symbol</td></tr>\n              <tr><td><code>offers.availability</code></td><td>Whether it can be bought now. Agents skip products they think are out of stock</td></tr>\n              <tr><td><code>sku</code>, <code>gtin</code>, <code>mpn</code></td><td>Let an agent match your product to the same item elsewhere when comparing</td></tr>\n              <tr><td><code>brand</code></td><td>Answers \"do you have anything from …?\" questions</td></tr>\n              <tr><td><code>aggregateRating</code></td><td>Lets agents weigh quality, but only include it if the reviews are on the page</td></tr>\n              <tr><td><code>hasMerchantReturnPolicy</code>, <code>shippingDetails</code></td><td>Answer the questions people ask before buying, without a trip to your policy pages</td></tr>\n            </tbody>\n          </table>\n          <h2>Products with options</h2>\n          <p>If a product comes in sizes or colors with different prices or stock, describe each variant. Schema.org's <code>ProductGroup</code> type groups variants under one parent, with <code>variesBy</code> naming what changes (such as size or color) and <code>hasVariant</code> listing each <code>Product</code> with its own offer. At minimum, make sure the price and availability on the page match the variant a visitor lands on.</p>\n          <h2>Beyond product pages</h2>\n          <ul>\n            <li><strong>Home page:</strong> <code>Organization</code> (or <code>LocalBusiness</code> with address and opening hours) with your name, logo, website, contact details and social profiles. AgentScore looks for structured data on your home page.</li>\n            <li><strong>Category pages:</strong> <code>BreadcrumbList</code> shows where a page sits in your catalog, and helps agents move between categories.</li>\n            <li><strong>Help pages:</strong> <code>FAQPage</code> for genuine question-and-answer content, such as shipping and returns questions.</li>\n            <li><strong>Software pricing:</strong> <code>SoftwareApplication</code> or <code>Product</code> with an <code>Offer</code> for each plan.</li>\n          </ul>\n          <h2>Rules that keep it trustworthy</h2>\n          <ul>\n            <li><strong>Match the page.</strong> Every value must match what a person sees. A price in the markup that differs from the page is worse than none: an agent may quote the wrong one, and search engines may ignore your markup.</li>\n            <li><strong>Keep it current.</strong> Generate it from the same data as the page, so prices and stock update together. Hand-written markup goes stale.</li>\n            <li><strong>Put it in the HTML.</strong> Markup added by JavaScript after the page loads is invisible to agents that don't run JavaScript. AgentScore flags this as a warning. See <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">why AI agents can't see JavaScript-only content</a>.</li>\n            <li><strong>Prefer JSON-LD.</strong> Microdata and RDFa work, but JSON-LD keeps the data in one tidy block that's easier to read and maintain.</li>\n            <li><strong>Don't mark up what isn't there.</strong> No ratings without visible reviews, no FAQ markup for questions that aren't on the page.</li>\n          </ul>\n          <h2>How to add it</h2>\n          <ol>\n            <li><strong>Check what you already have.</strong> Many e-commerce platforms and themes output basic product markup. View the page source and search for <code>application/ld+json</code>.</li>\n            <li><strong>Fill the gaps.</strong> Common missing pieces are <code>availability</code>, identifiers like GTIN, and return and shipping details. Platform apps and SEO plugins can add these, or your developer can extend the theme template.</li>\n            <li><strong>Validate.</strong> Paste the page URL into the <a href=\"https://validator.schema.org/\">Schema Markup Validator</a> to check the markup is valid, and Google's Rich Results Test to see which fields Google can use.</li>\n            <li><strong>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>.</strong> The structured data check passes when your home page has JSON-LD in its HTML, and warns if it only appears after JavaScript runs.</li>\n          </ol>\n          <p>Structured data is one part of being readable to agents. The other big one is making sure your content is in the HTML at all; see the guide to <a href=\"https://ghostagentlab.com/articles/javascript-content-ai-agents/\">JavaScript-only content</a>, or the full <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">agent readiness guide</a>.</p>",
      "date_published": "2026-10-08T12:05:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/javascript-content-ai-agents/",
      "url": "https://ghostagentlab.com/articles/javascript-content-ai-agents/",
      "title": "Why AI agents can't see your JavaScript-only content",
      "summary": "Many AI agents read only the HTML your server sends. How to check what they see, what usually goes missing, and how to fix it with server rendering.",
      "content_html": "<p>Open a modern website with JavaScript turned off and you'll often see a logo, a spinner and not much else. That's what many AI agents see too. If your prices, products or policies only appear after scripts run, a large share of AI traffic never sees them.</p>\n          <h2>How agents read a page</h2>\n          <p>When a browser loads a page, it does two things: it downloads the HTML your server sends, then runs the JavaScript, which may fetch more data and build much of what you see. Agents differ in how far they go:</p>\n          <table>\n            <thead><tr><th>Agent</th><th>Typically reads</th></tr></thead>\n            <tbody>\n              <tr><td>Training and AI search crawlers</td><td>The HTML only. Running JavaScript for billions of pages is slow and expensive, so most AI crawlers don't.</td></tr>\n              <tr><td>Assistant fetchers (fetching a page because someone asked)</td><td>Usually the HTML only, and they have a few seconds at most.</td></tr>\n              <tr><td>Browser agents</td><td>The full page, including JavaScript, but they're slow and impatient. Long loading states and content that appears late still trip them up.</td></tr>\n            </tbody>\n          </table>\n          <p>Googlebot does run JavaScript, which is why many sites that rely on it rank fine in Google search. That has hidden the problem: the same site can be well indexed by Google and nearly blank to AI assistants.</p>\n          <h2>How to check your own site</h2>\n          <ol>\n            <li><strong>View the source.</strong> In your browser, use View Page Source (not the developer tools inspector, which shows the page after JavaScript). Search for a product name or price you can see on the page. If it isn't in the source, it's added by JavaScript.</li>\n            <li><strong>Turn JavaScript off.</strong> In Chrome's developer tools, open the command menu, type \"Disable JavaScript\", and reload. What's left is roughly what an HTML-only agent sees.</li>\n            <li><strong>Fetch it from the command line.</strong> <code>curl -s https://yourdomain.com/your-page | grep -i \"your price\"</code> shows exactly what your server sends.</li>\n            <li><strong>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>.</strong> It compares the text in your HTML with the text after JavaScript runs. It passes when at least 70% of the text is already in the HTML, and warns or fails below that.</li>\n          </ol>\n          <h2>What usually goes missing</h2>\n          <ul>\n            <li><strong>Prices and stock</strong> loaded from an API after the page appears, so the HTML says \"Loading…\" or nothing at all.</li>\n            <li><strong>Whole pages in single-page apps</strong>, where the HTML is an empty <code>&lt;div id=\"root\"&gt;</code> and everything else is built in the browser.</li>\n            <li><strong>Product lists with infinite scroll</strong> that load more items only as a person scrolls, with no paginated links to follow.</li>\n            <li><strong>Content inside tabs and accordions</strong> fetched only when clicked, such as specifications, sizing and shipping details.</li>\n            <li><strong>Structured data</strong> injected by a tag manager or script, which agents reading the HTML never see. See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data AI shopping agents can read</a>.</li>\n            <li><strong>Reviews</strong> from third-party widgets that load in the browser.</li>\n          </ul>\n          <h2>How to fix it</h2>\n          <p>The aim is simple: the important content should be in the HTML your server sends. JavaScript can still make the page interactive afterwards.</p>\n          <h3>Server-side rendering (SSR)</h3>\n          <p>The server builds the full HTML for each request, then JavaScript takes over in the browser. Modern frameworks support this directly: Next.js and Remix for React, Nuxt for Vue, SvelteKit for Svelte, and Angular's server rendering. If your site is a single-page app built with one of these, turning on server rendering for key pages is usually the biggest single fix.</p>\n          <h3>Static generation</h3>\n          <p>Pages are built to HTML ahead of time, at deploy. It's ideal for content that changes rarely, such as marketing pages, docs, policies and blog posts. Astro, Eleventy, Hugo and the frameworks above can all do it, often alongside server rendering for pages that change often.</p>\n          <h3>Render the essentials, enhance the rest</h3>\n          <p>You don't need to render everything on the server. Prioritize what agents ask about: product name, price, availability, key specifications, shipping and returns, and plan pricing. Personalized recommendations, recently viewed items and chat widgets can stay in JavaScript.</p>\n          <h3>Pre-rendering as a stopgap</h3>\n          <p>Pre-rendering services run your JavaScript ahead of time and serve the finished HTML to bots. It can work as a short-term fix, but it means maintaining two versions of each page, and it depends on correctly recognizing every agent. Google describes serving bots differently (\"dynamic rendering\") as a workaround rather than a long-term solution. Server rendering is better.</p>\n          <h2>Smaller fixes that help</h2>\n          <ul>\n            <li>Replace infinite scroll with, or add, real paginated links (<code>?page=2</code>) agents can follow.</li>\n            <li>Put tab and accordion content in the HTML and hide it with CSS, instead of fetching it on click.</li>\n            <li>Move JSON-LD from your tag manager into the page template.</li>\n            <li>Include a summary of reviews (average rating and count) in the HTML, even if the full reviews load later.</li>\n            <li>Give the page a meaningful <code>&lt;title&gt;</code> and <code>&lt;h1&gt;</code> in the HTML, never only after JavaScript runs.</li>\n          </ul>\n          <p>Once your content is in the HTML, the next question is whether agents can find their way around it and use it. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>.</p>",
      "date_published": "2026-10-08T12:04:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Readability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/agent-friendly-buttons-forms/",
      "url": "https://ghostagentlab.com/articles/agent-friendly-buttons-forms/",
      "title": "Buttons, links and forms agents can use: a checklist",
      "summary": "Browser agents navigate by the names and roles of controls. Fix icon-only buttons, fake buttons, unlabelled fields, blocking pop-ups and hover-only menus.",
      "content_html": "<p>Browser agents don't see your page the way people do. They read its structure: what each button, link and form field is called, and what it does. A shopping-bag icon that a person recognizes instantly can be completely invisible to an agent. This checklist covers the fixes that matter most.</p>\n          <h2>How agents find their way around</h2>\n          <p>Most browser agents work from the page's <em>accessibility tree</em>, the same structured view of the page that screen readers use, often alongside a screenshot. In that tree, every element has a role (button, link, text box, heading) and a name (\"Add to cart\", \"Search\", \"Email address\"). The agent decides what to do by reading those names.</p>\n          <p>That's good news: almost everything that makes a site work for agents also makes it work for people using screen readers and keyboards. If you've invested in accessibility, you're most of the way there.</p>\n          <figure>\n            <div>\n              <div>\n              </div>\n              <div>\n                <p>What an agent reads</p>\n                <ul>\n                  <li><span><code>button</code> with no name (the search icon)</span></li>\n                  <li><span><code>button</code> with no name (the cart icon)</span></li>\n                  <li><span><code>heading</code> &ldquo;Trail runner&rdquo;</span></li>\n                  <li><span><code>text</code> &ldquo;$129&rdquo;</span></li>\n                  <li><span><code>generic</code>: the color swatches aren&rsquo;t controls at all</span></li>\n                  <li><span><code>button</code> &ldquo;Add to cart&rdquo;</span></li>\n                </ul>\n              </div>\n            </div>\n            <figcaption>The same product page two ways. Icon-only buttons have no name, and swatches built from plain boxes aren&rsquo;t controls, so the only thing an agent can use with confidence is the button with words on it.</figcaption>\n          </figure>\n          <h2>1. Give every button and link a name</h2>\n          <p>The most common failure is the icon-only control: a cart, search, menu or close button that shows an icon and has no text.</p>\n          <pre><code>&lt;!-- An agent sees: button, no name --&gt;\n&lt;button&gt;&lt;svg&gt;…&lt;/svg&gt;&lt;/button&gt;\n&lt;!-- An agent sees: button \"Cart, 2 items\" --&gt;\n&lt;button aria-label=\"Cart, 2 items\"&gt;&lt;svg aria-hidden=\"true\"&gt;…&lt;/svg&gt;&lt;/button&gt;</code></pre>\n          <ul>\n            <li>Use visible text where you can. Use <code>aria-label</code> where the design calls for an icon alone.</li>\n            <li>For a linked image, the image's <code>alt</code> text becomes the link's name.</li>\n            <li>Make names specific. Ten links called \"Learn more\" or \"Shop now\" tell an agent nothing about where each goes; \"Shop running shoes\" does.</li>\n          </ul>\n          <h2>2. Use real buttons and links</h2>\n          <p>A <code>&lt;div&gt;</code> with a click handler looks like a button to a person but isn't one to an agent: it has no role, can't be focused, and often isn't recognized as clickable at all.</p>\n          <pre><code>&lt;!-- Looks clickable, isn't recognized as clickable --&gt;\n&lt;div class=\"btn\" onclick=\"addToCart()\"&gt;Add to cart&lt;/div&gt;\n&lt;!-- Recognized as a button by every agent --&gt;\n&lt;button type=\"button\" onclick=\"addToCart()\"&gt;Add to cart&lt;/button&gt;</code></pre>\n          <ul>\n            <li>Use <code>&lt;a href&gt;</code> for anything that goes to another page, and <code>&lt;button&gt;</code> for anything that does something on this page.</li>\n            <li>If you truly can't change the element, add <code>role=\"button\"</code> and <code>tabindex=\"0\"</code>, and handle the Enter and Space keys.</li>\n            <li>Make links real URLs. A link with <code>href=\"#\"</code> and a script gives an agent nowhere to go if the script doesn't run.</li>\n          </ul>\n          <h2>3. Label every form field</h2>\n          <p>Agents fill in forms by matching what they've been asked to enter (\"my email\", \"ship to 10001\") to the field's label. Placeholder text is a weak substitute: it disappears once someone types, and not every agent reads it.</p>\n          <pre><code>&lt;label for=\"email\"&gt;Email address&lt;/label&gt;\n&lt;input id=\"email\" name=\"email\" type=\"email\" autocomplete=\"email\" required&gt;</code></pre>\n          <ul>\n            <li>Connect each <code>&lt;label&gt;</code> to its field with <code>for</code> and <code>id</code>, or wrap the field in the label.</li>\n            <li>Use the right <code>type</code> (<code>email</code>, <code>tel</code>, <code>number</code>) and <code>autocomplete</code> values (<code>given-name</code>, <code>postal-code</code>, <code>address-line1</code>). They tell an agent exactly what each field wants.</li>\n            <li>Prefer native <code>&lt;select&gt;</code>, checkboxes and radio buttons over custom-built dropdowns. Custom size and color pickers are a frequent place agents get stuck.</li>\n            <li>Show errors as text next to the field and link them with <code>aria-describedby</code>, so an agent knows what to fix. A red border alone isn't enough.</li>\n          </ul>\n          <h2>4. Get pop-ups and banners out of the way</h2>\n          <p>Cookie banners, newsletter pop-ups and region selectors that cover the page on arrival are one of the biggest blockers. An agent has to work out how to dismiss them before it can do anything else, and if it can't, it's stuck.</p>\n          <ul>\n            <li>Close and accept controls must be real buttons with clear names: \"Accept all\", \"Reject all\", \"Close\".</li>\n            <li>Keep consent banners small, at the edge of the screen, rather than covering the content.</li>\n            <li>Don't show sign-up pop-ups on arrival. If you must, wait until someone has engaged with the page.</li>\n          </ul>\n          <h2>5. Make menus work without hover</h2>\n          <p>Agents click; they don't hover. A menu that only opens when the mouse rests over it may never open for an agent, hiding every category inside it.</p>\n          <ul>\n            <li>Make the top-level menu item a button that opens the menu on click, with <code>aria-expanded</code> showing whether it's open.</li>\n            <li>Make sure the top-level items are also real links to category pages, so there's always a route in.</li>\n          </ul>\n          <h2>6. Mark up the page's landmarks</h2>\n          <p>Landmarks tell an agent where the main content is, and where the navigation is, so it can skip straight to what matters.</p>\n          <pre><code>&lt;header&gt;…&lt;/header&gt;\n&lt;nav aria-label=\"Main\"&gt;…&lt;/nav&gt;\n&lt;main&gt;…&lt;/main&gt;\n&lt;footer&gt;…&lt;/footer&gt;</code></pre>\n          <p>Use one <code>&lt;h1&gt;</code> per page and headings in order, so the page has an outline an agent can follow.</p>\n          <h2>How to test it</h2>\n          <ol>\n            <li><strong>Tab through the page.</strong> Use only the keyboard to reach the menu, search, a product, its options and the cart. Anywhere you get stuck, an agent probably does too.</li>\n            <li><strong>Look at the accessibility tree.</strong> In Chrome's developer tools, the Accessibility pane shows each element's role and name. Look for buttons with no name.</li>\n            <li><strong>Run an accessibility checker</strong>, such as Lighthouse in Chrome or axe, for a quick list of unnamed controls and unlabelled fields.</li>\n            <li><strong>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>.</strong> It renders your home page in a browser and checks named buttons and links, real controls, form labels, landmarks and pop-ups that cover the page. To see where a real AI agent gets stuck on a journey, run a <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agent</a> test.</li>\n          </ol>\n          <div>\n            <p><strong>Test the journeys, not just the page.</strong> A home page can pass every check while the size picker on a product page or the checkout button stops agents cold. <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a> run your key journeys, such as search, add to cart and sign-up, on a schedule, and alert you when one breaks.</p>\n          </div>\n          <p>For the bigger picture, see <a href=\"https://ghostagentlab.com/articles/agent-readiness/\">What is agent readiness?</a></p>",
      "date_published": "2026-10-08T12:03:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Navigability"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/agent-ready-checkout/",
      "url": "https://ghostagentlab.com/articles/agent-ready-checkout/",
      "title": "How to make checkout work for AI shopping agents",
      "summary": "A step-by-step guide to the shopping journey for AI agents: finding products, choosing options, adding to cart, checkout and the payment handoff.",
      "content_html": "<p>When a person asks an AI assistant to buy something, the agent has to do everything a shopper does: find the product, choose the right options, add it to the cart, and get through checkout. Each step is a place it can get stuck, and a stuck agent is a lost sale. This guide walks the journey step by step.</p>\n          <h2>How agents shop today</h2>\n          <p>Agents buy in two ways, and you'll want to support both:</p>\n          <ul>\n            <li><strong>Through your website.</strong> A browser agent uses your site the way a person would, and typically hands control back to the person to sign in or pay. This works today on any site, as long as the agent can get through.</li>\n            <li><strong>Through an agentic commerce protocol.</strong> The assistant completes the purchase through a direct integration, without driving your web pages. Examples include the Agentic Commerce Protocol used by ChatGPT and Google's Universal Commerce Protocol. Support usually comes through your commerce platform or payment provider.</li>\n          </ul>\n          <p>Protocols are growing fast, but for now most agents still shop through the website. That's the journey this guide covers.</p>\n          <figure>\n            <ol>\n              <li>\n                <p>Find the product</p>\n                <p>Hover-only menus, and search that only works with scripts</p>\n              </li>\n              <li>\n                <p>Choose options <span>Most common failure</span></p>\n                <p>Custom size and color swatches with no names or roles</p>\n              </li>\n              <li>\n                <p>Add to cart</p>\n                <p>A button covered by a banner, or no text confirming the add</p>\n              </li>\n              <li>\n                <p>The cart</p>\n                <p>Quantity and remove buttons shown only as icons</p>\n              </li>\n              <li>\n                <p>Checkout</p>\n                <p>Forced accounts, unlabelled fields and a CAPTCHA on every order</p>\n              </li>\n              <li>\n                <p>Payment <span>Handed to the person</span></p>\n                <p>The agent hands back to the person; a cart that empties during the handoff loses the sale</p>\n              </li>\n            </ol>\n            <figcaption>The six steps an agent works through for a shopper, and what most often stops it at each. Choosing options is where agents most often fail on otherwise good sites.</figcaption>\n          </figure>\n          <h2>Step 1: Find the product</h2>\n          <ul>\n            <li><strong>Search works without JavaScript tricks.</strong> A real search form with a labelled field and a submit button, whose results have their own URL (like <code>/search?q=decaf</code>), lets agents search directly.</li>\n            <li><strong>Categories are real links.</strong> Category and product links should be ordinary <code>&lt;a href&gt;</code> links in the HTML, reachable without hover menus. See <a href=\"https://ghostagentlab.com/articles/agent-friendly-buttons-forms/\">buttons, links and forms agents can use</a>.</li>\n            <li><strong>Product pages describe the product.</strong> Name, price, stock and key specifications in the HTML, with structured data. See <a href=\"https://ghostagentlab.com/articles/structured-data-ai-shopping-agents/\">product data agents can read</a>.</li>\n          </ul>\n          <h2>Step 2: Choose options</h2>\n          <p>This is where agents most often fail on otherwise good sites. Size, color and quantity pickers are frequently custom-built swatches that look fine to people but have no names or roles agents can read.</p>\n          <ul>\n            <li>Use native <code>&lt;select&gt;</code> elements or radio buttons, or give custom swatches a role and a name, for example a radio group with each option labelled \"Size 10\" or \"Color: Navy\".</li>\n            <li>Show unavailable options as disabled, with text saying why (\"Out of stock\"), not just a faded color.</li>\n            <li>Update price and stock text on the page when an option changes, so the agent can confirm it picked correctly.</li>\n            <li>Give each variant its own URL where possible (<code>?variant=…</code>), so an agent can link straight to the exact item.</li>\n          </ul>\n          <h2>Step 3: Add to cart</h2>\n          <ul>\n            <li>The add-to-cart control is a real <code>&lt;button&gt;</code> with clear text, and nothing covers it, such as a cookie banner or chat widget.</li>\n            <li>Confirm the add in text: \"Added to cart: Ethiopia Yirgacheffe, 340 g\" with a link to the cart. An animation alone leaves the agent unsure whether it worked.</li>\n            <li>If a cart drawer slides out, make sure it has a labelled close button and a real link to the cart page.</li>\n            <li>Skip upsell pop-ups at this step, or make them easy to dismiss with a labelled button.</li>\n          </ul>\n          <h2>Step 4: The cart</h2>\n          <ul>\n            <li>The cart lives at a stable URL, such as <code>/cart</code>, that agents can open directly.</li>\n            <li>Line items, quantities, subtotal and any discounts are shown as text.</li>\n            <li>Quantity controls and remove buttons have names (\"Remove Ethiopia Yirgacheffe\"), not just \"×\" or a trash icon.</li>\n            <li>The checkout button is a clearly named button or link, visible without scrolling past recommendations.</li>\n          </ul>\n          <h2>Step 5: Checkout</h2>\n          <ul>\n            <li><strong>Offer guest checkout.</strong> Forcing account creation is a dead end for many agents. Offer sign-in as an option, not a gate.</li>\n            <li><strong>Label every field and use autocomplete.</strong> <code>autocomplete=\"shipping street-address\"</code>, <code>postal-code</code>, <code>email</code> and <code>tel</code> tell an agent exactly where each detail goes.</li>\n            <li><strong>Make shipping options readable.</strong> Each option's name, price and delivery estimate as text, with a real radio button to choose it.</li>\n            <li><strong>Write errors as text.</strong> \"Enter a 5-digit ZIP code\", next to the field, rather than a red outline.</li>\n            <li><strong>Show the full total before payment</strong>, including shipping and tax, so the agent can confirm it with the person before they pay.</li>\n            <li><strong>Avoid a CAPTCHA on every checkout.</strong> Use risk-based checks that only challenge suspicious orders. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a>.</li>\n          </ul>\n          <h2>Step 6: Payment</h2>\n          <p>Most agents stop at payment and hand back to the person, and that's how it should be. Make the handoff smooth: a clear payment step, wallet options like Apple Pay, Google Pay or Shop Pay that a person can approve in one tap, and a cart that survives the handoff instead of emptying after a short timeout.</p>\n          <p>To let assistants complete purchases on their own, ask your commerce platform or payment provider which agentic commerce protocols it supports. AgentScore checks whether your site advertises one.</p>\n          <h2>How to test your checkout</h2>\n          <ol>\n            <li><strong>Walk it with the keyboard only.</strong> From the home page to the payment step, without a mouse. Anywhere you get stuck, an agent probably does too.</li>\n            <li><strong>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>.</strong> For online stores, it checks that AI agents can open your product, cart and checkout pages, find and press the add-to-cart button, and fill in labelled checkout fields, and that there's a guest checkout. It doesn't run a full journey itself.</li>\n            <li><strong>Monitor it with <a href=\"https://ghostagentlab.com/ghost-agent/\">Ghost Agents</a>.</strong> They run the whole journey on a schedule with test details you provide, and alert you when it breaks. They stop when they reach a payment form, never enter card details, and refuse buttons that place orders, so testing never creates a real order.</li>\n          </ol>\n          <div>\n            <p><strong>Retest after every change.</strong> Checkout breaks for agents in the same ways it breaks for people: a theme update, a new app, a redesigned size picker. The difference is that agents don't email support. They just buy elsewhere.</p>\n          </div>",
      "date_published": "2026-10-08T12:02:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Task completion"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/measure-ai-agent-traffic/",
      "url": "https://ghostagentlab.com/articles/measure-ai-agent-traffic/",
      "title": "Why your analytics can't see AI agents, and how to measure agent traffic",
      "summary": "JavaScript analytics tags miss most AI agents. How to measure agent traffic from server and CDN logs, and what to track: who's visiting, who's real, and where they're blocked.",
      "content_html": "<p>Ask most businesses how many AI agents visited their site last month and they can't say. It's not that the agents aren't coming. It's that the tools most teams use to count visitors were built in a way that can't see them. Here's why, and how to measure agent traffic properly.</p>\n          <h2>Why analytics misses agents</h2>\n          <p>Tools like Google Analytics count a visit when a small JavaScript tag runs in the visitor's browser and reports back. That works for people. It fails for AI agents for three reasons:</p>\n          <ul>\n            <li><strong>Most agents never run the tag.</strong> AI crawlers and assistant fetchers download the HTML and leave without running JavaScript, so the tag never fires.</li>\n            <li><strong>Known bots are filtered out on purpose.</strong> Analytics tools exclude known bot traffic so it doesn't distort your human numbers. That's right for marketing reports, but it hides exactly the visits you want to see.</li>\n            <li><strong>Browser agents blend in.</strong> Agents that do run JavaScript often look like ordinary browsers, so they're counted as people, mixed in with everyone else.</li>\n          </ul>\n          <p>You may see AI in your analytics as referral traffic: people who click a link in a ChatGPT or Perplexity answer and arrive on your site. That's valuable, but it's the result of agent visits, not the visits themselves. It doesn't show which agents read your pages, which they couldn't, or what they saw.</p>\n          <h2>Where agent visits do show up</h2>\n          <p>Every request to your site, from a person or an agent, passes through your web server or CDN, and they can log it whether or not any JavaScript runs. Each log line typically includes:</p>\n          <table>\n            <thead><tr><th>Field</th><th>What it tells you</th></tr></thead>\n            <tbody>\n              <tr><td>User agent</td><td>Which agent it claims to be</td></tr>\n              <tr><td>IP address</td><td>Where it came from, which you need to verify it's real</td></tr>\n              <tr><td>Path</td><td>Which page or file it asked for</td></tr>\n              <tr><td>Status code</td><td>Whether it got the page (200), was blocked (403), rate limited (429) or hit a missing page (404)</td></tr>\n              <tr><td>Time and bytes</td><td>When it came, and how much content it received</td></tr>\n            </tbody>\n          </table>\n          <h2>Getting your logs</h2>\n          <ul>\n            <li><strong>Cloudflare:</strong> Logpush sends HTTP request logs to storage or another service. Availability depends on your plan.</li>\n            <li><strong>Vercel:</strong> log drains stream request logs to an endpoint you choose.</li>\n            <li><strong>Fastly, Akamai and Amazon CloudFront</strong> all offer access logs or real-time log streaming.</li>\n            <li><strong>Your own servers:</strong> Nginx and Apache write access logs by default.</li>\n            <li><strong>Hosted platforms</strong> such as Shopify don't usually give you raw request logs, which makes agent traffic much harder to see. Ask your platform what bot and crawler reporting it offers.</li>\n          </ul>\n          <h2>A quick look with the command line</h2>\n          <p>If you have an Nginx or Apache access log in the standard \"combined\" format, this counts requests from some well-known AI agents:</p>\n          <pre><code>awk -F'\"' '{print $6}' access.log \\\n  | grep -oiE 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Applebot|Amazonbot|meta-external[a-z]+|Bytespider|CCBot' \\\n  | sort | uniq -c | sort -rn</code></pre>\n          <p>It's a useful first look, but it has real limits: it trusts the user agent, so impostors count as the real thing; it only knows the names you list; and it says nothing about whether the agents succeeded.</p>\n          <h2>What to measure</h2>\n          <ol>\n            <li><strong>Who's visiting.</strong> Requests by agent and by kind: training crawlers, AI search, assistant fetchers and browser agents. The mix matters more than the total. Assistant fetchers are people asking about you right now.</li>\n            <li><strong>Who's real.</strong> Verify each agent against its operator's published IP ranges or reverse DNS, and separate verified, spoofed and \"can't tell\". See <a href=\"https://ghostagentlab.com/articles/verify-ai-crawlers/\">how to tell if an AI crawler is real</a>. Without this, scrapers inflate the numbers.</li>\n            <li><strong>What they're reading.</strong> The pages agents request most. Are they reaching products, pricing and policies, or stuck on the home page and old URLs?</li>\n            <li><strong>Whether they're blocked.</strong> The share of agent requests that get 403, 429 or challenge responses. A rise usually means a bot protection change caught good agents. See <a href=\"https://ghostagentlab.com/articles/bot-protection-ai-agents/\">bot protection and CAPTCHAs</a>.</li>\n            <li><strong>Errors.</strong> 404s for agents often point to old URLs that AI models learned and still request. Redirect them.</li>\n            <li><strong>Trends.</strong> Week over week, by agent. New agents appear often, and a sudden drop from one usually means something on your side changed.</li>\n          </ol>\n          <h2>Privacy</h2>\n          <p>Logs contain IP addresses, which are personal data for human visitors. You only need the full detail for bots and agents. Hash or drop IP addresses for human traffic, keep agent records separately, and set a retention period that matches your privacy policy.</p>\n          <div>\n            <p><strong>Ghost Agent Labs does this for you.</strong> Connect Cloudflare Logpush, a Vercel log drain, or any JSON log feed. Every request is matched against a registry of more than 1,500 known agents, crawlers are verified against published IP ranges and reverse DNS, and you get a dashboard of who's visiting, which agents are real, what they read, and where they're blocked. Human visits are only counted, never stored. <a href=\"https://app.ghostagentlab.com/signup\">Start free</a>.</p>\n          </div>\n          <p>Measuring traffic tells you who's coming. To find out whether they succeed, test the journeys that matter. See <a href=\"https://ghostagentlab.com/articles/agent-ready-checkout/\">how to make checkout work for AI shopping agents</a>.</p>",
      "date_published": "2026-10-08T12:01:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Measurement"
      ]
    },
    {
      "id": "https://ghostagentlab.com/articles/llms-txt/",
      "url": "https://ghostagentlab.com/articles/llms-txt/",
      "title": "How to write an llms.txt file",
      "summary": "What llms.txt is, what to put in it, a complete example for an online store, and the mistakes that make it useless to AI agents.",
      "content_html": "<p>An llms.txt file is a short Markdown guide to your website, written for AI. It tells an agent what your business does and where the pages that matter are, so it doesn't have to work it out from menus, pop-ups and marketing copy. It takes about half an hour to write, and AgentScore checks for it.</p>\n          <h2>What llms.txt is</h2>\n          <p><a href=\"https://llmstxt.org/\">llms.txt</a> is a proposed standard, first published by Jeremy Howard of Answer.AI in 2024. The idea is simple: put a plain-text Markdown file at <code>/llms.txt</code> on your domain, the same way <code>/robots.txt</code> lives at the root. Where robots.txt tells crawlers what they may visit, llms.txt tells AI models and agents what is worth reading.</p>\n          <p>It's useful because AI agents have limited attention. A typical page is mostly navigation, scripts, tracking and layout. An agent asked \"what's the return policy at this store?\" has to dig through all of that. A good llms.txt hands it the answer's location in one request.</p>\n          <div>\n            <p><strong>An honest caveat:</strong> llms.txt is a proposal, not something every AI company has committed to read. Some coding assistants and documentation tools use it today, and it's cheap to add, but treat it as one part of agent readiness alongside clear pages, structured data and a sitemap, not a replacement for them.</p>\n          </div>\n          <h2>The format</h2>\n          <p>The file is ordinary Markdown with a light structure, in this order:</p>\n          <ol>\n            <li><strong>An H1 with your name.</strong> This is the only required part.</li>\n            <li><strong>A blockquote summary</strong>: one or two sentences on what you do and who for.</li>\n            <li><strong>Optional paragraphs</strong> with anything an agent should know up front, such as where you ship or what you don't sell.</li>\n            <li><strong>H2 sections of links</strong>, each a Markdown list item: <code>- [Page name](URL): what it's for</code>.</li>\n            <li><strong>An \"Optional\" section</strong> for links an agent can skip if it's short on space.</li>\n          </ol>\n          <h2>A complete example</h2>\n          <p>Here's an llms.txt for a fictional online store:</p>\n          <pre><code># Northwind Coffee\n&gt; Northwind Coffee roasts and sells specialty coffee beans, ground coffee and\n&gt; brewing equipment online, shipping to the US and Canada.\nOrders over $40 ship free in the US. We don't sell gift cards or wholesale online;\nwholesale enquiries go through the contact page.\n## Shop\n- [All coffee](https://northwind.example/collections/coffee): Every bean we sell, with roast level, origin and price\n- [Subscriptions](https://northwind.example/subscriptions): Recurring deliveries every 2, 4 or 6 weeks; pause or cancel any time\n- [Brewing equipment](https://northwind.example/collections/equipment): Grinders, kettles and pour-over kits\n## Help and policies\n- [Shipping](https://northwind.example/pages/shipping): Rates, delivery times and countries we ship to\n- [Returns](https://northwind.example/pages/returns): 30-day returns on unopened items and equipment\n- [FAQ](https://northwind.example/pages/faq): Grind sizes, freshness, subscriptions and account questions\n- [Contact](https://northwind.example/pages/contact): Email and live chat hours\n## Optional\n- [Our story](https://northwind.example/pages/about): Who we are and how we source our beans\n- [Brewing guides](https://northwind.example/blog/guides): Step-by-step guides for each brewing method</code></pre>\n          <p>You can also see <a href=\"https://ghostagentlab.com/llms.txt\">our own llms.txt</a>.</p>\n          <h2>What to include</h2>\n          <p>Think about the questions people ask agents about businesses like yours, and link to the page that answers each one.</p>\n          <table>\n            <thead><tr><th>Type of site</th><th>Link to</th></tr></thead>\n            <tbody>\n              <tr><td>Online store</td><td>Main categories, bestsellers, shipping, returns, sizing, FAQ, contact</td></tr>\n              <tr><td>Software company</td><td>Product overview, pricing, sign-up, docs, API reference, security, status page</td></tr>\n              <tr><td>Local or service business</td><td>Services and prices, booking, locations and hours, service area, contact</td></tr>\n              <tr><td>Publisher</td><td>Sections, subscription options, licensing and permissions, editorial policy</td></tr>\n            </tbody>\n          </table>\n          <p>If you offer agents a direct way in, such as an API or an MCP server, list it with a link to its documentation. AgentScore looks for an MCP server mentioned in llms.txt as one of the signals that a site supports agents directly.</p>\n          <h2>Common mistakes</h2>\n          <ul>\n            <li><strong>Serving HTML.</strong> Some site builders return your home page or a 404 page for any unknown URL. Open <code>/llms.txt</code> in a browser and make sure you see plain text, not a web page.</li>\n            <li><strong>Writing marketing copy.</strong> \"The world's most loved coffee experience\" tells an agent nothing. Say what you sell, where, and on what terms.</li>\n            <li><strong>Linking to everything.</strong> It's a guide, not a sitemap. Twenty to forty well-described links beats five hundred bare ones. Your sitemap.xml already lists every page.</li>\n            <li><strong>Leaving out descriptions.</strong> The text after each link is what lets an agent pick the right page without opening them all.</li>\n            <li><strong>Letting it go stale.</strong> Dead links and old prices are worse than no file. Review it whenever you change your navigation or policies.</li>\n            <li><strong>Linking to pages agents can't read.</strong> If a linked page only shows content after JavaScript runs, or sits behind a CAPTCHA, the link won't help. AgentScore checks for both.</li>\n          </ul>\n          <h2>How to publish it</h2>\n          <ol>\n            <li>Write the file in any text editor and save it as <code>llms.txt</code>.</li>\n            <li>Upload it so it's served at the root of your domain: <code>https://yourdomain.com/llms.txt</code>. On most hosts this means putting it in the same folder as robots.txt. On Shopify, WordPress and other platforms, apps and plugins can serve it for you.</li>\n            <li>Check it's served as plain text with a 200 status, and that your bot protection doesn't challenge AI agents that request it.</li>\n            <li>Run <a href=\"https://ghostagentlab.com/agentscore/\">AgentScore</a>. The \"llms.txt guide for AI\" check will pass once the file is live.</li>\n          </ol>\n          <h2>What about llms-full.txt?</h2>\n          <p>Some sites, mostly documentation sites, also publish <code>/llms-full.txt</code>: the full text of their key pages in one Markdown file, so an agent can load everything at once. It's a community convention rather than part of the proposal. It's worth it for developer docs; for most stores and marketing sites a good llms.txt is enough.</p>",
      "date_published": "2026-10-08T12:00:00+00:00",
      "authors": [
        {
          "name": "Ghost Agent Labs"
        }
      ],
      "tags": [
        "Article",
        "Access"
      ]
    }
  ]
}
