A sitemap is a list of the pages on your site that you want found, in a format machines read without effort. Search engines have used them for about twenty years. AI agents and AI search crawlers use them too, to find pages your navigation hides and to discover new products quickly. It's one of the cheapest readiness wins there is.
What a sitemap does for AI agents
An agent arriving at your site has two ways to find things: follow links from the home page, or read a list someone has prepared. Following links is slow and fragile. Menus built with JavaScript, products buried four clicks deep and pages that only appear in search results can all be missed.
A sitemap solves that. It's an XML file, usually at /sitemap.xml, that lists your important URLs. Crawlers from AI search engines use it to decide what to fetch. Agents doing a task can use it to locate a product or policy page directly. And it gives every system the same, complete picture of what you publish.
It doesn't replace good links. A page that's only in your sitemap is still a page most visitors and many agents never reach. Treat the sitemap as the safety net under your navigation, not instead of it. Our guide to linking key pages from the home page covers the other half.
What AgentScore checks
The "Valid sitemap" check, in the Readability category, looks for your sitemap the way crawlers do:
- It reads
Sitemap:lines in your robots.txt. - It then tries
/sitemap.xmland/sitemap_index.xml. - For each, it checks the file loads with a 200 status and is a real sitemap: XML with a
<urlset>or<sitemapindex>element.
It passes when it finds one, and reports how many entries it holds and whether it's listed in robots.txt. It fails if none of those places has a valid sitemap. A home page or a "not found" page served at /sitemap.xml doesn't count, which catches a common problem with some site builders.
The sitemap also feeds other checks. AgentScore uses it to find a product or pricing page when the home page doesn't link to one. If the sitemap is an index, it reads product sitemaps first, since most platforms name them that way. That's why a missing sitemap can leave other checks marked "not tested": without a product page to look at, there's nothing to check prices or product data on.
What a good sitemap looks like
The format is defined at sitemaps.org. A minimal one for a small store:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://northwind.example/</loc>
<lastmod>2026-10-01</lastmod>
</url>
<url>
<loc>https://northwind.example/products/ethiopia-yirgacheffe</loc>
<lastmod>2026-10-07</lastmod>
</url>
<url>
<loc>https://northwind.example/pages/shipping</loc>
<lastmod>2026-08-14</lastmod>
</url>
</urlset>
Larger sites split it up. Each sitemap file can hold up to 50,000 URLs, and a sitemap index lists the individual files:
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap><loc>https://northwind.example/sitemap_products_1.xml</loc></sitemap>
<sitemap><loc>https://northwind.example/sitemap_collections_1.xml</loc></sitemap>
<sitemap><loc>https://northwind.example/sitemap_pages_1.xml</loc></sitemap>
</sitemapindex>
Then tell crawlers where it is, with one line in robots.txt:
Sitemap: https://northwind.example/sitemap.xml
What to include, and what to leave out
| Include | Leave out |
|---|---|
| Product and category pages you want people to find | Cart, checkout and account pages |
| Pricing, plans and sign-up pages | Internal search results and filtered URLs |
| Shipping, returns, FAQ and contact pages | Pages that redirect elsewhere |
| Guides, articles and other evergreen content | Pages marked noindex, or blocked in robots.txt |
| The canonical URL of each page | Duplicates with tracking parameters or session IDs |
The rule of thumb: list the pages you'd be happy for an AI assistant to send a customer to, each at its one true address. If you include a page in the sitemap but block it in robots.txt, you're sending crawlers mixed signals.
Keep it accurate
- Generate it automatically. Shopify, WooCommerce, Magento, BigCommerce, WordPress SEO plugins and most modern frameworks produce a sitemap for you. A hand-made sitemap goes stale the week after it's written.
- Make
lastmodhonest. It should change when the content does, such as a price or stock change, not every time the file is rebuilt. Google's documentation says it useslastmodonly when it's consistently accurate, and other crawlers are likely to treat it similarly. - Don't bother with priority and changefreq. They're in the protocol, but Google says it ignores them. Your time is better spent on accurate dates.
- Remove dead pages. Discontinued products that return 404, or redirect to a category, shouldn't stay listed.
- Keep it reachable. Make sure bot protection doesn't challenge crawlers asking for the sitemap. See our guide to bot protection and CAPTCHAs.
Sitemap, llms.txt or both?
Both, because they do different jobs. A sitemap is complete: every page you want found, with no explanation. An llms.txt file is selective: a short, described guide to the pages that matter most, written for AI. A crawler indexing your catalog wants the sitemap. An assistant trying to answer "what's their returns window?" benefits from llms.txt pointing straight at the returns page.
Quick check: open yourdomain.com/robots.txt and look for a Sitemap: line. Then open the URL it gives. You should see XML listing your pages, not a web page. If either step fails, that's your first fix.
Next steps
- Confirm your platform generates a sitemap, and that it loads.
- Add a
Sitemap:line to robots.txt if it's missing. - Spot-check that your bestselling products and your policy pages are listed.
- Run AgentScore. The "Valid sitemap" check will show what it found and how many entries it counted.
For how the sitemap fits with the other things agents need, see the agent readiness guide.