# Canonical URLs and duplicate pages for AI agents

> How to stop AI agents reading the same page at many URLs: rel=canonical, parameter and variant URLs, http, https and www redirects, and clean sitemaps.

Published October 9, 2026 by Ghost Agent Labs · Readability · https://ghostagentlab.com/articles/canonical-urls-ai-agents/

## Key takeaways

- Stores show the same product at many web addresses, and naming one main address helps AI agents read and repeat the right version.
- Without it, agents may treat one product as several, quote an old price, or send shoppers to a page with a stale campaign tag.
- Set a canonical link on every page, redirect old http and www versions in one step, and list only main addresses in your sitemap.

Most online stores show the same product at many different addresses: with tracking tags, with a color selected, under two categories, on http and https. People don't notice. AI agents do, because each address can look like a different page. Canonical URLs tell them which one is the real one.

## Why duplicate URLs confuse AI agents

When an AI agent reads your site, or when an AI search service builds the index an assistant draws on, it meets your pages by URL. If the same jacket appears at five URLs, it may:

- treat them as five products, and compare your jacket with itself;
- quote a price or stock level from an old copy it read weeks ago;
- link a shopper to a version with a stale campaign tag, an odd filter or a variant they didn't ask for;
- spend its limited time on your site reading copies instead of finding new products.

The fix is not to make duplicates disappear, which is often impossible, but to say clearly which URL is the main one, and to make every other signal agree.

## Where duplicates come from

| Source | Example |
| --- | --- |
| Tracking parameters | `/products/rain-jacket?utm_source=newsletter` |
| Sort and filter parameters | `/jackets?sort=price&color=blue` |
| Variant parameters | `/products/rain-jacket?variant=4417` |
| Category paths | `/mens/jackets/rain-jacket` and `/sale/rain-jacket` |
| Protocol and host | `http://` and `https://`, with and without `www.` |
| Trailing slashes and case | `/Rain-Jacket/` and `/rain-jacket` |

## Set a canonical URL on every page

A canonical link is one line in the page's `<head>` that names the preferred URL for the content:

```
<link rel="canonical" href="https://northwind.example/products/rain-jacket">
```

Put it on every indexable page, including on the canonical page itself (a self-referencing canonical). Rules that keep it reliable:

- **Use the full, absolute URL.** Include `https://` and the host, not just the path.
- **Put it in the HTML.** A canonical tag added by JavaScript may never be seen by agents and crawlers that read the raw page.
- **One per page.** Two different canonical tags on a page is a common theme or app conflict. Search engines may ignore both.
- **Point to a page that works.** The canonical should return 200, not redirect, 404 or carry a `noindex`.
- **Don't canonicalize everything to the home page.** Each distinct page should point at itself, or at its true duplicate.

Search engines treat `rel="canonical"` as a strong hint rather than a command, and the same is likely true of AI services that build on them. That's why the other signals below need to agree with it.

## Parameter URLs

Tracking parameters like `utm_source` and click IDs never change the content, so the page should always name the clean URL as canonical. Sort orders and session IDs are the same.

Filters are a judgment call. A filtered listing such as "blue jackets" may be worth keeping as its own page if people search for it. If so, give it a clean path (`/jackets/blue`), its own title and a self-referencing canonical. Most combinations of filters, though, should point back to the unfiltered category. For more on filters agents can use, see [site search and filters for AI agents](https://ghostagentlab.com/articles/site-search-filters-ai-agents/).

## Product variants

Variants are where stores most often go wrong. There are two reasonable patterns:

- **One page for the product.** Every variant URL (`?variant=4417`, `?color=blue`) has a canonical pointing to the main product URL. This suits products where variants differ only in size or color and share a description. Make sure the main page lists every variant's price and availability, ideally in `ProductGroup` markup, so the agent doesn't lose them.
- **A page per variant.** Each variant has its own URL with a self-referencing canonical, its own title ("Rain Jacket, Men's, Blue") and its own price. This suits variants that are really different products, with different specs or prices.

What doesn't work is a mix: variant URLs that canonicalize to the main page, while the main page only shows the default variant's price and stock. An agent is then told the blue jacket is the same page as the red one, but can't find the blue one's details. See [size, color and variant pickers AI agents can use](https://ghostagentlab.com/articles/variant-pickers-ai-agents/).

## http, https and www: redirect, don't just canonicalize

For duplicates that are purely technical, a canonical tag isn't enough. Send a permanent (301 or 308) redirect so that every visitor, person or agent, lands on one version:

- `http://` redirects to `https://`.
- Pick `www.northwind.example` or `northwind.example`, and redirect the other.
- Pick trailing slash or no trailing slash, and redirect the other.
- Redirect old product URLs to their replacements after a migration, not to the home page.

Redirect in a single hop. Chains (http to https, then to www, then to a trailing slash) slow every request down, and some clients stop following after a few hops. Avoid redirects that depend on JavaScript or a meta refresh; agents that don't run scripts get stuck on the first page. And make sure region or language redirects don't trap agents, as covered in [international sites for AI agents](https://ghostagentlab.com/articles/international-sites-ai-agents/).

## Make every signal agree

Agents and crawlers look at more than the canonical tag. Each of these should use the same, canonical URL:

- **Your XML sitemap.** List only canonical URLs that return 200. No parameter URLs, no redirecting URLs, no http versions. See [XML sitemaps for AI agents](https://ghostagentlab.com/articles/xml-sitemaps-ai-agents/).
- **Internal links.** Navigation, product grids and breadcrumbs should link to the canonical URL directly, not to a version that redirects or carries parameters.
- **Structured data.** The `url` in your Product and Offer JSON-LD should match the canonical.
- **`og:url` and feeds.** Social tags and product feeds should use the same address.
- **hreflang tags**, if you have them, should point to canonical URLs of each language version.

## What AgentScore checks

[AgentScore](https://ghostagentlab.com/agentscore/) doesn't have a dedicated canonical check. The "Valid sitemap" check confirms you have a sitemap it can read (found through robots.txt, `/sitemap.xml` or `/sitemap_index.xml`) and counts its entries; it doesn't judge whether each URL is canonical. AgentScore also uses your sitemap and home page links to find a product page to test, so a sitemap full of redirecting or parameter URLs makes that harder for any agent.

## A quick audit

1. Open a product page with `?utm_source=test` on the end. View the source and check the canonical points to the clean URL.
2. Do the same for a variant URL and a filtered category. Is the result the pattern you intended?
3. Type `http://` and the non-preferred host into a browser. Do they each redirect, in one hop, to the right page?
4. Open your sitemap and spot-check ten URLs: do they all return 200, and match their own canonical?
5. Check your main navigation links for parameters or redirects.

Canonical URLs are housekeeping, but they decide which version of your page agents read and repeat. For the bigger picture, see the [agent readiness guide](https://ghostagentlab.com/articles/agent-readiness/).

## Sources and further reading

- [How to specify a canonical with rel="canonical" and other methods](https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls) (Google Search Central)
- [Redirects and Google Search](https://developers.google.com/search/docs/crawling-indexing/301-redirects) (Google Search Central)
- [RFC 6596: The Canonical Link Relation](https://www.rfc-editor.org/rfc/rfc6596.html) (IETF)
