Sign in Start free

Faceted Navigation Crawl Control

Filter combinations multiply: five facets with six values each produce more URLs than you have products. The obvious fix - canonical tags pointing back to the bare category - fails because canonicals are a hint, not a directive, and Google still has to crawl every URL to read them. The budget is burnt before the signal is even seen. You need to stop the URLs being discovered, not just deduplicate them after the fact.

CATEGORY
Ecommerce SEO
FORMAT
faceted-navigation-crawl-control.md
STEPS
6
PRICE
Free - no account
WHEN TO REACH FOR THIS

Use when a category with a few hundred products has generated tens of thousands of crawlable filter URLs and crawl stats show Googlebot spending its budget on them.

The skill file

faceted-navigation-crawl-control.md
---
name: faceted-navigation-crawl-control
description: Use when a category with a few hundred products has generated tens of thousands of crawlable filter URLs and crawl stats show Googlebot spending its budget on them.
---

# Faceted Navigation Crawl Control

Filter combinations multiply: five facets with six values each produce more URLs than you have products. The obvious fix - canonical tags pointing back to the bare category - fails because canonicals are a hint, not a directive, and Google still has to crawl every URL to read them. The budget is burnt before the signal is even seen. You need to stop the URLs being discovered, not just deduplicate them after the fact.

## What you need first

- a full list of facet parameters and how they combine in URLs
- server log data or Search Console crawl stats showing which parameter URLs get crawled
- search demand data per facet value (color, size, brand, price)

## Method

1. Export 30 days of server logs and count Googlebot hits by URL pattern. Anything where parameter URLs exceed 30 percent of total crawl is an active problem, not a theoretical one.
2. Split facets into three tiers: indexable (facet values with real search demand, usually brand, material, and a handful of attribute terms), crawlable but noindex, and never-linked.
3. For the never-linked tier, render the filter controls as POST forms or JavaScript-driven state that produces no anchor tag - Google cannot follow what is not an href. This is the only reliable containment.
4. Give indexable facets clean static paths, not query strings: /boots/waterproof/ rather than /boots?feature=waterproof. Static paths get treated as real pages and support their own copy and internal links.
5. Block the remaining parameter patterns in robots.txt only once nothing indexed depends on them, and confirm with a site: check first - blocking an indexed URL freezes it in the index with no way to remove it.
6. Cap multi-select depth: allow one facet selected, noindex two, and refuse to generate a link at three. Recheck logs at 30 days to confirm parameter crawl share has fallen.

## What this produces

A tiered facet policy document plus the template changes that stop non-indexable combinations being linked at all.

## Where this goes wrong

- relying on rel=canonical alone, which still costs a crawl per URL and gets ignored when the page content differs
- blocking parameters in robots.txt while those URLs are still indexed, which strips the canonical signal and strands them
- indexing price and sort facets, which change constantly and produce thin near-duplicate pages with zero search demand

---

From the QuQi skill library - https://www.quqi.io/skills/faceted-navigation-crawl-control
Free to download · no account, no email

What you need first

  • a full list of facet parameters and how they combine in URLs
  • server log data or Search Console crawl stats showing which parameter URLs get crawled
  • search demand data per facet value (color, size, brand, price)

Method

  1. 01 Export 30 days of server logs and count Googlebot hits by URL pattern. Anything where parameter URLs exceed 30 percent of total crawl is an active problem, not a theoretical one.
  2. 02 Split facets into three tiers: indexable (facet values with real search demand, usually brand, material, and a handful of attribute terms), crawlable but noindex, and never-linked.
  3. 03 For the never-linked tier, render the filter controls as POST forms or JavaScript-driven state that produces no anchor tag - Google cannot follow what is not an href. This is the only reliable containment.
  4. 04 Give indexable facets clean static paths, not query strings: /boots/waterproof/ rather than /boots?feature=waterproof. Static paths get treated as real pages and support their own copy and internal links.
  5. 05 Block the remaining parameter patterns in robots.txt only once nothing indexed depends on them, and confirm with a site: check first - blocking an indexed URL freezes it in the index with no way to remove it.
  6. 06 Cap multi-select depth: allow one facet selected, noindex two, and refuse to generate a link at three. Recheck logs at 30 days to confirm parameter crawl share has fallen.

What this produces

A tiered facet policy document plus the template changes that stop non-indexable combinations being linked at all.

Where this goes wrong

  • relying on rel=canonical alone, which still costs a crawl per URL and gets ignored when the page content differs
  • blocking parameters in robots.txt while those URLs are still indexed, which strips the canonical signal and strands them
  • indexing price and sort facets, which change constantly and produce thin near-duplicate pages with zero search demand

Use this skill in your own AI

The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.

Claude Code Save it as ~/.claude/skills/faceted-navigation-crawl-control/SKILL.md and Claude loads it on its own when what you are doing matches the trigger line. Put it in .claude/skills inside a project instead if the whole team should have it.
Claude Upload the file in the skills section of your settings. Once it is there it applies itself in any conversation where the trigger fits, so you do not have to remember it exists.
ChatGPT There is no skills format to install into, so paste the file contents into a Project instruction or a Custom GPT instead. It then applies to every chat in that project rather than only the one you paste it into.
Anything else Paste the markdown into the chat before your question. It works in any assistant, it just has to be pasted again each time.

Questions about this skill

When is crawl control the right method rather than canonical tags on filter URLs?

When crawl stats or server logs show parameter URLs taking a large share of Googlebot hits. The obvious alternative, a canonical pointing back to the bare category, still costs one crawl per URL before the hint is read, and gets ignored where the filtered content differs. Reach for containment when the problem is discovery volume rather than duplication in the index.

What do I need before tiering the facets?

A full list of facet parameters and how they combine, thirty days of logs or Search Console crawl stats, and search demand per facet value. Without demand data you tier facets by type and throw away the two brands or materials people genuinely search for. Without logs you cannot tell whether parameter crawl is thirty per cent of the budget or three.

What does the facet policy consist of, and which part does the work?

A three tier facet policy and the template changes behind it. The policy is the record, but the working part is the template change that stops the never linked tier emitting an anchor at all, because Google cannot follow what is not an href. Moving the indexable tier to static paths is the other half that ships as code.

What is the mistake that most often ruins this?

Blocking parameter patterns in robots.txt while those URLs are still indexed. The crawler can then no longer fetch them to see a canonical or a noindex, so they freeze in the index with no clean route out. Run the site: check first and block only once nothing indexed depends on the pattern, or you have made removal harder than the bloat was.

More in Ecommerce SEO