What scanning 600 stores taught us about our own scanner
To write our study of online stores, we ran AgentScore on 600 real websites in one day. The study taught us a lot about stores. It taught us even more about our own scanner: three bugs that tests on our own fixture sites never caught, and one run whose results looked perfect because almost nothing had been tested. Here's what we found, what we changed, and what it means for your score.
Why run the scanner at scale
Every AgentScore check is tested against small, carefully built example sites. That proves the check does what we meant. It doesn't prove we meant the right thing for the thousands of ways real stores are built. Scanning 600 stores, from many countries, platforms and sizes, was the first time AgentScore met that variety all at once, and the first time we read its results in aggregate, where a check that's wrong for many stores stands out.
Bug 1: AgentScore only spoke English
Several checks work by reading what a button or link says: finding the add-to-cart button, recognising the button that closes a cookie banner, spotting the link to the cart. All of them only knew English words. On a German store, "In den Warenkorb" wasn't recognised as an add-to-cart button, and "Alle akzeptieren" didn't count as a way to close the banner.
In our first run, 61 stores whose pages weren't in English were reported as having no add-to-cart button at all. When we rescanned eight of them with the fix, AgentScore found the button on six. The other two turned out to be a different problem (see bug 3).
There was a subtler issue underneath. The usual way to match whole words in code only understands the 26 letters of the English alphabet, so even a correctly translated phrase could fail to match at the edges of words in Russian or with accented letters. Japanese, Chinese and Korean don't put spaces between words at all.
What we changed: the words AgentScore looks for now live in one list covering 19 languages, including German, French, Spanish, Italian, Dutch, the Nordic languages, Polish, Turkish, Russian, Ukrainian, Japanese, Chinese and Korean, and matching works in every alphabet. It also changed which sites AgentScore recognises as stores, because finding the cart link is part of that: the same sample produced 230 stores instead of 219.
Bug 2: Real buttons reported as fake ones
AgentScore checks that your add-to-cart button is a real button, because agents look for buttons and links, not for text that happens to be clickable. To find the button, it looked for the smallest element whose text matched. Many shop themes write their buttons like this:
<button type="submit">
<span>Add to cart</span>
</button>
The smallest element containing "Add to cart" is the <span>, so AgentScore reported "a plain <span>, not a real button", on a perfectly good button. That's a false alarm, and an expensive one: in the study's second run, 57% of the stores where we found the button were flagged. After the fix, the true figure is 24%.
What we changed: when the matching text sits inside a real button or link, AgentScore now judges that button or link. A <div> styled to look like a button is still caught, because that really is a problem for agents.
Bug 3: Sold-out products counted against stores
AgentScore checks one product page per store. When that product was sold out, its buy button had been replaced by a disabled "Sold out" button, "売り切れ" on a Japanese store or "Agotado" on a Spanish one, and AgentScore reported that it couldn't find an add-to-cart button. That says nothing about the store's buttons, only about its stock.
What we changed: AgentScore now recognises a sold-out product, from its button in any of the 19 languages or from the product data saying it's out of stock. Then it does what a shopper would: it tries up to two other products. If they're all sold out, the check is marked as not tested, which never costs points. In the study, AgentScore moved past a sold-out product on 4 stores and skipped the check on 5 where everything it tried was sold out.
The run that looked perfect
One of our runs came back looking wonderful. Navigation scored 100%. Not one store had an unnamed button. No pop-ups blocked anything.
It was wrong. After an update to how AgentScore opens pages in its browser, the browser on our scanning machine couldn't complete secure connections, and it couldn't load a single page. Every check that needs a real browser was skipped, on 245 of 245 stores. Skipped checks don't cost points, by design, so the scores went up instead of down.
We caught it because the numbers were too good to be true, not because anything failed. That's not good enough, so we added a safeguard: our analysis now refuses to produce results if more than 10% of stores couldn't be opened in a browser. We fixed the scanning machine's setup, confirmed one scan by hand, and ran all 600 sites again.
Your own scans are protected the same way. When AgentScore can't open your pages in a browser, your report says so and marks those checks as not tested. It never presents a skipped check as a pass. See how to read your AgentScore report.
What this means for your score
- Stores outside English-speaking markets get a fairer score, and some will see it rise, as buttons, banners and cart links are now recognised in their language.
- Stores whose buttons wrap their text in another element no longer get a false "not a real button" finding.
- A sold-out product no longer costs you points.
All three changes are in AgentScore's current scoring version, 2026.10-v5. If you scanned your site before, run AgentScore again to see your updated score.
What we took away
- Test against the real world, not only your examples. Each bug passed every test we had, because our test sites were built the way we expected sites to be built.
- Be suspicious of good news. The broken run would have made a better headline than the real one. A result that's surprisingly good deserves the same scrutiny as one that's surprisingly bad.
- Publish the limits. The study states what AgentScore can't see, such as checkouts that need items in the cart and requests that don't come from AI companies' verified addresses. Saying so is what makes the rest of the numbers worth trusting.
Curious how AgentScore opens and reads your pages? How AgentScore renders pages explains the browser side, and How AgentScore works covers the scoring.