Articles

How to test your key journeys with AI agents

Usually fixed by: Developer

Technical checks tell you whether AI agents should be able to use your site. The only way to know whether they can is to send one through and watch. This guide covers which journeys to test, how to test them by hand with the assistants your customers already use, how to automate it, and how to keep testing safe.

Why checks aren't enough

A site can pass every check for labels, named buttons and structured data and still lose agents at a custom size picker, a cookie banner that reappears on the cart page, or a checkout step that needs an account. Those failures only show up when an agent actually tries to finish the job. That's why task completion carries the most weight in agent readiness, and why it can only be measured by running real agents.

Step 1: Pick the journeys that make money

Don't test everything. Start with the two or three journeys where an agent failure costs you a customer:

Type of siteJourneys to test first
Online storeSearch for a product, choose options and add to cart, check out as a guest up to payment
Software companyFind pricing, start a trial or sign up, book a demo
Service businessFind prices and availability, book an appointment, request a quote
Any siteFind a policy (returns, shipping, cancellation), use the contact form

Write each one the way a customer would ask an assistant: "Find a waterproof jacket in medium under $150 and add it to the cart." Specific goals give clearer results than "buy something".

Step 2: Test by hand with consumer assistants

The quickest start costs nothing. Several consumer AI assistants can now browse and act on websites, usually in an "agent" or browsing mode. Use the ones your customers are likely to use.

  1. Give it the goal and your site. "On northwind.example, find a 12-cup coffee maker and add it to the cart. Don't buy anything."
  2. Watch, don't help. Most agent modes show what the agent is doing. Note every hesitation, wrong click and retry, not just whether it finished.
  3. Ask it what went wrong. If it fails, ask what it was trying to do and what it couldn't find. The answer often names the exact control.
  4. Repeat it. Agents don't behave the same way every time. Run each journey at least three times before drawing conclusions.

Manual testing is good for discovery: it shows you how agents experience your site and builds a shared picture for the team. It doesn't scale, and it won't tell you when something breaks next Tuesday.

Step 3: Automate it with Ghost Agents

Ghost Agents are AI agents that run your key journeys for you. In the Ghost Agent Labs app, you set up a test for a site you've added:

  • Goal. Start from a common journey (checkout, product search, sign-up or contact form) or describe your own in plain words.
  • Start page. Where the agent begins. It can only visit pages on your site.
  • When it passes. A check on the final page the agent reaches: it reached the payment step, the page shows certain text, or the address contains a certain path. These are checked directly, not by asking an AI model whether it succeeded.
  • Test data. A test email address and any other values the journey needs.
  • Schedule and repeats. Run it by hand, daily or hourly, with several runs per check.

Each run records every step: the page the agent saw, what it did and its reasoning. When a run fails, the replay opens on the step where it got stuck, and you can compare it with the last run that passed. Failing tests can alert your team by email or Slack.

What to measure

  • Pass rate. Because agents vary, one run proves little. Run each check several times and look at the share that passed. Ghost Agents report a journey as passing, flaky (some runs passed) or failing (none did).
  • Where it fails. The step where failures cluster is your fix list: the size picker, the cookie banner, the account wall.
  • Steps taken. A journey that passes in 12 steps one week and 25 the next has become harder, even if it still passes.
  • Change over time. A journey that passed yesterday and fails today usually means something on your site changed: a theme update, a new app, a new pop-up.

Flaky is a finding, not noise. If an agent completes a journey two times in three, a real customer's assistant fails one time in three. Treat a flaky result as a problem to fix, and look at the failed runs to see what was different.

Keep testing safe

Testing with agents means letting software click buttons on your live site. Set some rules:

  • No real purchases. Stop at the payment step. Ghost Agents stop when they reach a payment form, never enter card details, and refuse buttons that place orders, pay, or delete or cancel accounts. When testing by hand, tell the assistant not to buy, and stay in control at payment.
  • Test data only. Use a dedicated test email address, such as one on a domain you control, and dummy names and phone numbers. Never use real card or bank details, or a real customer's details.
  • Watch your side effects. Test sign-ups create accounts, test leads land in your CRM and test bookings take slots. Use recognizable test values (a "test+" email or the name "Ghost Test") so your team can filter them out, and avoid submitting booking forms that hold real inventory.
  • Tell the people who need to know. Let whoever runs bot protection and analytics know tests are coming. Ghost Agents identify themselves with GhostAgent/1.0 in the user agent and sign their requests with Web Bot Auth, so they can be recognized and filtered from reports.
  • Consider staging for risky journeys. If a journey can't be tested on the live site without real consequences, test it on a staging copy that matches production.

Turn results into fixes

A failed run is only useful if someone acts on it. For each failure, record the journey, the step, what the agent was trying to do and a screenshot, then route it to whoever owns that part of the site: the developer for an unlabelled button, the platform admin for forced account creation, marketing for the pop-up. The replay usually tells you which.

Most fixes are the same ones covered elsewhere in these guides: buttons and forms agents can use, cookie banners and pop-ups, checkout and sign-up and booking forms.

Where to start this week

  1. Run AgentScore on your home page to clear the obvious blockers first.
  2. Pick your single most valuable journey and try it by hand with an AI assistant, three times.
  3. Set it up as a Ghost Agent test with a daily schedule and a pass condition on the final page, so you hear about it the day it breaks.
← All articles Test your site with AgentScore →

Can AI agents finish the job on your site?

Ghost Agents work through your journeys from a plain-English goal, with a step-by-step replay of where they got stuck. Included in every plan, even Free.