Reading your AgentScore report: what to fix first
An AgentScore report gives you a number, a sentence and a list of checks. The number tells you where you stand. The list tells you what to do about it, as long as you read it in the right order. This guide walks through each part of the report, in the free scan and in the Ghost Agent Labs app, and shows how to turn it into a short list of fixes with an owner for each.
Start at the top: score, band and verdict
The score is out of 100. It falls into one of four bands, and the verdict under it says what the band means:
| Score | Band | Verdict |
|---|---|---|
| 80 and up | Good | AI agents can use this site well. |
| 60 to 79 | Fair | AI agents can mostly use this site, but some things get in their way. |
| 40 to 59 | Weak | AI agents will struggle on this site. |
| Below 40 | Poor | Most AI agents will fail on this site. |
If any check failed, the verdict ends with Biggest issue: the summary of the failed check that carries the most weight. It's the single best place to start. If nothing failed, there's no biggest issue, and your work is in the warnings.
One thing to know before you share the number: the free scan doesn't run an AI agent through a task, so Task completion shows as "n/a" and the score is worked out over Access, Readability and Navigability. A 72 means 72% of the points that could be tested. How AgentScore works has the arithmetic.
Then the categories
Each category gets a bar and its points, such as Access 20/25. The app shows the same thing as a share of the category's points, so Access 20/25 reads as 80/100. Use the bars to see where points are being lost, not to rank your work: a low Navigability bar with one failure in it can matter less than a single Access failure.
Access comes first for a reason. If AI agents are blocked by your bot protection or robots.txt, they never see the pages the other two categories are about.
What each status means
| Status | What it means | Counts toward the score |
|---|---|---|
| Passed | Agents will be fine here. | Full weight |
| Warning | It works, but agents will stumble, or a newer standard is missing. | Half weight |
| Failed | This stops or seriously misleads agents. | Nothing |
| Not tested | The check couldn't run on your site. | Left out entirely |
Checks are listed failures first, then warnings, then passes, then not tested. Each one has a one-line summary of what was found, often with an example, such as how many buttons have no name or which assistant was blocked.
Don't skip "not tested"
A "not tested" check never costs you points, and its summary always says why it was skipped. The reasons fall into a few groups:
- It doesn't apply. Guest checkout, checkout fields, product data and agentic commerce are for stores. If the scan found no products or cart, they're skipped.
- It needs a full cart. The scanner never adds anything to a cart, so a checkout that only opens with items in it can't be inspected.
- A page couldn't be rendered. If the browser couldn't load a page, the checks that depend on it are skipped and say so. If it happens on every scan, ask a developer to check whether that page loads for a first-time desktop visitor.
- It isn't in the free scan. "An AI agent completes a real task" is always not tested there.
Read these anyway. If you run a store and the report says no product page was found from the home page or sitemap, that's worth knowing in itself: agents look for products the same way the scanner does.
Fixes, severity and score impact
Every warning and failure comes with a fix: the specific change to make. In the free report, you see the findings straight away and unlock the fixes with your email address. In the app, every fix is shown.
The app's Findings page adds three things that make the list easier to work through:
- Severity. A failure on a check with weight 4 or more is Critical, weight 2 or 3 is High, and weight 1 is Medium. A warning is Medium on a check with weight 3 or more, otherwise Low. A missing llms.txt is labelled Opportunity: worth doing, but a new standard, weighted lightly.
- Score impact. How many points your overall score would gain if that finding were fixed. On a store where every Access check was tested, a failed Bot protection lets AI agents through is worth about 7 points; a missing llms.txt is worth about 1. The numbers depend on which checks could be tested on your site, so read them from your own report.
- Why it matters and who usually fixes it. Open "Why it matters and how to fix it" under any finding for a plain-English reason, the fix, and the role that usually owns it: Developer, Bot protection or CDN admin, SEO or content team, or E-commerce platform admin.
The Readiness page shows the top four as Biggest wins. The Findings page lets you filter by severity, and Export CSV downloads every finding with its severity, score impact, summary, why it matters, fix, who usually fixes it and the scan date, ready for a ticketing tool.
How to decide what to fix first
- Fix Access failures first. Blocked assistants, challenges on arrival and CAPTCHAs at the checkout stop agents before anything else matters. These usually sit with whoever runs your CDN or bot protection, and are often a settings change rather than a project. Our guide to bot protection and AI agents covers the common causes.
- Then the remaining Critical and High findings, by score impact. Content that only appears after JavaScript, a banner covering the add-to-cart button, and product pages with no price data are typical.
- Group by owner, not by category. Sort the CSV by "Usually fixed by" and send each person their own short list. One ticket per owner moves faster than one long list for everyone.
- Pick up cheap warnings along the way. Image alt text, page landmarks and a missing meta description are small changes that often ship with other work.
- Leave opportunities until the basics pass. llms.txt, MCP and agentic commerce are worth doing, and score lightly for a reason: they help most once agents can already get in and read your pages.
- Rescan after each batch. In the app you can rescan whenever you like and watch the trend. The free page reuses a domain's result for 24 hours.
A rule of thumb: if a finding means agents can't get in or can't see your prices, it's this sprint. If it means they'll find it a little harder, it's this quarter. If it's a new standard, it's on the roadmap. Our 90-day plan lays that out week by week.
What the report can't tell you
A high score means your site removes the common obstacles. It doesn't prove that an AI agent can find a product, choose a size and reach the payment step. That's what Ghost Agents test in the app: real AI agents run your key journeys, stop at the payment form, and give you a step-by-step replay of where they got stuck. Use the report to clear the path, and journey tests to prove it's clear.