Faceted Navigation Crawl Control
Filter combinations multiply: five facets with six values each produce more URLs than you have products. The obvious fix - canonical tags pointing back to the bare category - fails because canonicals are a hint, not a directive, and Google still has to crawl every URL to read them. The budget is burnt before the signal is even seen. You need to stop the URLs being discovered, not just deduplicate them after the fact.
Use when a category with a few hundred products has generated tens of thousands of crawlable filter URLs and crawl stats show Googlebot spending its budget on them.
The skill file
What you need first
- a full list of facet parameters and how they combine in URLs
- server log data or Search Console crawl stats showing which parameter URLs get crawled
- search demand data per facet value (color, size, brand, price)
Method
- 01 Export 30 days of server logs and count Googlebot hits by URL pattern. Anything where parameter URLs exceed 30 percent of total crawl is an active problem, not a theoretical one.
- 02 Split facets into three tiers: indexable (facet values with real search demand, usually brand, material, and a handful of attribute terms), crawlable but noindex, and never-linked.
- 03 For the never-linked tier, render the filter controls as POST forms or JavaScript-driven state that produces no anchor tag - Google cannot follow what is not an href. This is the only reliable containment.
- 04 Give indexable facets clean static paths, not query strings: /boots/waterproof/ rather than /boots?feature=waterproof. Static paths get treated as real pages and support their own copy and internal links.
- 05 Block the remaining parameter patterns in robots.txt only once nothing indexed depends on them, and confirm with a site: check first - blocking an indexed URL freezes it in the index with no way to remove it.
- 06 Cap multi-select depth: allow one facet selected, noindex two, and refuse to generate a link at three. Recheck logs at 30 days to confirm parameter crawl share has fallen.
What this produces
A tiered facet policy document plus the template changes that stop non-indexable combinations being linked at all.
Where this goes wrong
- relying on rel=canonical alone, which still costs a crawl per URL and gets ignored when the page content differs
- blocking parameters in robots.txt while those URLs are still indexed, which strips the canonical signal and strands them
- indexing price and sort facets, which change constantly and produce thin near-duplicate pages with zero search demand
Use this skill in your own AI
The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.
Questions about this skill
When is crawl control the right method rather than canonical tags on filter URLs?
When crawl stats or server logs show parameter URLs taking a large share of Googlebot hits. The obvious alternative, a canonical pointing back to the bare category, still costs one crawl per URL before the hint is read, and gets ignored where the filtered content differs. Reach for containment when the problem is discovery volume rather than duplication in the index.
What do I need before tiering the facets?
A full list of facet parameters and how they combine, thirty days of logs or Search Console crawl stats, and search demand per facet value. Without demand data you tier facets by type and throw away the two brands or materials people genuinely search for. Without logs you cannot tell whether parameter crawl is thirty per cent of the budget or three.
What does the facet policy consist of, and which part does the work?
A three tier facet policy and the template changes behind it. The policy is the record, but the working part is the template change that stops the never linked tier emitting an anchor at all, because Google cannot follow what is not an href. Moving the indexable tier to static paths is the other half that ships as code.
What is the mistake that most often ruins this?
Blocking parameter patterns in robots.txt while those URLs are still indexed. The crawler can then no longer fetch them to see a canonical or a noindex, so they freeze in the index with no clean route out. Run the site: check first and block only once nothing indexed depends on the pattern, or you have made removal harder than the bloat was.
More in Ecommerce SEO
Category Page Copy That Ranks
Use when category pages sit in positions 8 to 20 for their head term and the page is a bare grid...
Product Descriptions At Scale
Use when you have thousands of SKUs on manufacturer-supplied copy and product pages get impressi...
Out Of Stock And Discontinued Products
Use when products go out of stock or get discontinued and you need to decide what happens to the...
Product Schema That Earns Listings
Use when product pages are valid in the rich results test but show no price, rating, or availabi...