Blog

Ghost Agents: testing your site with real AI agents

A readiness score tells you what might trip an AI agent up. It can't tell you whether one actually gets from your home page to checkout. Ghost Agents do. They're real AI agents that work through your key journeys the way a customer's assistant would, on a schedule, and show you every step.

Why test with a real agent

Checks for labelled buttons, structured data and bot protection catch the common problems. But journeys fail in ways no rule predicts: a size picker the agent can't operate, a discount pop-up that appears on the second page, a checkout button that only shows after scrolling. People get past these without noticing. Agents often don't.

This is what our team has done for human visitors for years with Ghost Inspector: prove the journeys that make you money still work, and find out when a release breaks one. Ghost Agents bring the same idea to AI agents.

How a test works

You write a goal in plain English, the way a customer might ask an AI assistant. To make it quick, there are ready-made templates for the journeys that matter most:

TemplateWhat the agent tries to do
CheckoutFind a popular product, add it to the cart, and check out as a guest up to the payment step
Product searchUse site search to find a product a customer might ask for, and open its page
Sign-upCreate an account with a test email address, as far as the site allows without email confirmation
Contact formFind the contact or support form and fill it in with a short test question, without submitting it

On each run, the agent opens your site in a fresh browser and repeats a simple loop. It looks at the page as an agent sees it: the address, the visible text, and the buttons, links and fields with their names. It picks one action, such as click, type, choose an option, scroll, go back or finish. Then it acts, and looks again.

How you know it passed

An agent saying "done" isn't proof. So you choose pass conditions that are checked against the final page the agent reaches, without asking an AI model:

  • The page address contains some text, such as /checkout.
  • The page shows some text, such as "Thanks for signing up".
  • The agent reached the payment step. This is the one to use for checkout tests.
  • The agent reports the goal is done. This is used only if you set no other condition.

AI agents don't behave the same way every time, so each check runs the journey several times (three is a good default) and reports a pass rate. A test where every run passes is passing. One where none pass is failing. One where results disagree is flaky: agents can do the journey, but not reliably, which for a customer's assistant can mean the same as not at all.

Tests run when you click "Run now", every day, or every hour. When one fails, you can be alerted by email or Slack.

Replays: see exactly where it got stuck

Every run is recorded. The replay opens on the step where things went wrong, with a screenshot of what the agent saw, the action it took, and its reasoning, which often says in plain words what it couldn't find or press. If a journey used to work, you can compare the failed run with the last passing one to see what changed.

That turns "agents can't check out" into something a developer can act on: "the cookie banner covers the checkout button and has no labelled close button".

Safe by design

A test agent that clicks around a live store has to be careful. Ghost Agents follow hard rules that the model can't talk its way past:

  • They stop as soon as a payment form appears, and never enter card details. Reaching that point is how a checkout test passes.
  • They refuse buttons that place orders, pay, or delete or cancel an account.
  • They type only the test data you give the test, such as a test email address.
  • They stay on the site being tested.
  • Each run stops after 40 steps or five minutes, whichever comes first.

They also identify themselves. Every request carries GhostAgent/1.0 (+https://ghostagentlab.com/ghost-agent) at the end of a normal Chrome user agent and is signed with Web Bot Auth, so your CDN can verify it's really us. We explain signing in Why we sign every request our agents send, and the Ghost Agent page covers how to allow or opt out.

What Ghost Agents don't do yet

We'd rather you knew the limits up front:

  • They don't feed AgentScore yet. Task completion is the AgentScore category these tests will fill. Until they're connected, the score covers Access, Readability and Navigability. See How AgentScore works.
  • One agent per run. Every agent gets the same instructions and tools, so results from different AI models can be compared, but running several side by side in one check isn't available yet.
  • No release triggers yet. Tests run on demand or on a schedule, not automatically after each deploy.
  • No payments. By design, a checkout test proves an agent can reach payment. It never completes one.

Where to start

  1. Run AgentScore and fix anything that keeps agents out, such as blocked assistants or a challenge page. A Ghost Agent can't test a journey it can't start.
  2. In the app, add your site and create a Checkout test (or Sign-up, if you sell software). Run it once by hand.
  3. Watch the replay, even if it passed. You'll learn how agents read your pages.
  4. Set it to run daily and turn on alerts, so you hear about the release that breaks agent checkout before your customers' assistants do.

For help choosing journeys and writing good goals, read How to test your key journeys with AI agents.

← All blog Test your site with AgentScore →

Can AI agents use your site?

Get your free AgentScore in under a minute. No sign-up needed.