Most online stores show the same product at many different addresses: with tracking tags, with a color selected, under two categories, on http and https. People don't notice. AI agents do, because each address can look like a different page. Canonical URLs tell them which one is the real one.
Why duplicate URLs confuse AI agents
When an AI agent reads your site, or when an AI search service builds the index an assistant draws on, it meets your pages by URL. If the same jacket appears at five URLs, it may:
- treat them as five products, and compare your jacket with itself;
- quote a price or stock level from an old copy it read weeks ago;
- link a shopper to a version with a stale campaign tag, an odd filter or a variant they didn't ask for;
- spend its limited time on your site reading copies instead of finding new products.
The fix is not to make duplicates disappear, which is often impossible, but to say clearly which URL is the main one, and to make every other signal agree.
Where duplicates come from
| Source | Example |
|---|---|
| Tracking parameters | /products/rain-jacket?utm_source=newsletter |
| Sort and filter parameters | /jackets?sort=price&color=blue |
| Variant parameters | /products/rain-jacket?variant=4417 |
| Category paths | /mens/jackets/rain-jacket and /sale/rain-jacket |
| Protocol and host | http:// and https://, with and without www. |
| Trailing slashes and case | /Rain-Jacket/ and /rain-jacket |
Set a canonical URL on every page
A canonical link is one line in the page's <head> that names the preferred URL for the content:
<link rel="canonical" href="https://northwind.example/products/rain-jacket">
Put it on every indexable page, including on the canonical page itself (a self-referencing canonical). Rules that keep it reliable:
- Use the full, absolute URL. Include
https://and the host, not just the path. - Put it in the HTML. A canonical tag added by JavaScript may never be seen by agents and crawlers that read the raw page.
- One per page. Two different canonical tags on a page is a common theme or app conflict. Search engines may ignore both.
- Point to a page that works. The canonical should return 200, not redirect, 404 or carry a
noindex. - Don't canonicalize everything to the home page. Each distinct page should point at itself, or at its true duplicate.
Search engines treat rel="canonical" as a strong hint rather than a command, and the same is likely true of AI services that build on them. That's why the other signals below need to agree with it.
Parameter URLs
Tracking parameters like utm_source and click IDs never change the content, so the page should always name the clean URL as canonical. Sort orders and session IDs are the same.
Filters are a judgment call. A filtered listing such as "blue jackets" may be worth keeping as its own page if people search for it. If so, give it a clean path (/jackets/blue), its own title and a self-referencing canonical. Most combinations of filters, though, should point back to the unfiltered category. For more on filters agents can use, see site search and filters for AI agents.
Product variants
Variants are where stores most often go wrong. There are two reasonable patterns:
- One page for the product. Every variant URL (
?variant=4417,?color=blue) has a canonical pointing to the main product URL. This suits products where variants differ only in size or color and share a description. Make sure the main page lists every variant's price and availability, ideally inProductGroupmarkup, so the agent doesn't lose them. - A page per variant. Each variant has its own URL with a self-referencing canonical, its own title ("Rain Jacket, Men's, Blue") and its own price. This suits variants that are really different products, with different specs or prices.
What doesn't work is a mix: variant URLs that canonicalize to the main page, while the main page only shows the default variant's price and stock. An agent is then told the blue jacket is the same page as the red one, but can't find the blue one's details. See size, color and variant pickers AI agents can use.
http, https and www: redirect, don't just canonicalize
For duplicates that are purely technical, a canonical tag isn't enough. Send a permanent (301 or 308) redirect so that every visitor, person or agent, lands on one version:
http://redirects tohttps://.- Pick
www.northwind.exampleornorthwind.example, and redirect the other. - Pick trailing slash or no trailing slash, and redirect the other.
- Redirect old product URLs to their replacements after a migration, not to the home page.
Redirect in a single hop. Chains (http to https, then to www, then to a trailing slash) slow every request down, and some clients stop following after a few hops. Avoid redirects that depend on JavaScript or a meta refresh; agents that don't run scripts get stuck on the first page. And make sure region or language redirects don't trap agents, as covered in international sites for AI agents.
Make every signal agree
Agents and crawlers look at more than the canonical tag. Each of these should use the same, canonical URL:
- Your XML sitemap. List only canonical URLs that return 200. No parameter URLs, no redirecting URLs, no http versions. See XML sitemaps for AI agents.
- Internal links. Navigation, product grids and breadcrumbs should link to the canonical URL directly, not to a version that redirects or carries parameters.
- Structured data. The
urlin your Product and Offer JSON-LD should match the canonical. og:urland feeds. Social tags and product feeds should use the same address.- hreflang tags, if you have them, should point to canonical URLs of each language version.
What AgentScore checks
AgentScore doesn't have a dedicated canonical check. The "Valid sitemap" check confirms you have a sitemap it can read (found through robots.txt, /sitemap.xml or /sitemap_index.xml) and counts its entries; it doesn't judge whether each URL is canonical. AgentScore also uses your sitemap and home page links to find a product page to test, so a sitemap full of redirecting or parameter URLs makes that harder for any agent.
A quick audit
- Open a product page with
?utm_source=teston the end. View the source and check the canonical points to the clean URL. - Do the same for a variant URL and a filtered category. Is the result the pattern you intended?
- Type
http://and the non-preferred host into a browser. Do they each redirect, in one hop, to the right page? - Open your sitemap and spot-check ten URLs: do they all return 200, and match their own canonical?
- Check your main navigation links for parameters or redirects.
Canonical URLs are housekeeping, but they decide which version of your page agents read and repeat. For the bigger picture, see the agent readiness guide.