Sign in Start free

Crawl budget audit

Search engines allocate a finite amount of crawling to any site. On a catalog or a site with faceted navigation, that budget is routinely spent on near-duplicate URLs while the pages you care about wait. This skill finds where the budget is going and reclaims it.

CATEGORY
Technical SEO
FORMAT
crawl-budget-audit.md
STEPS
6
PRICE
Free - no account
WHEN TO REACH FOR THIS

Use when a large site has pages that stay unindexed for weeks, or when log files show crawlers spending their time on parameters, filters and pagination instead of the pages that matter.

The skill file

crawl-budget-audit.md
---
name: crawl-budget-audit
description: Use when a large site has pages that stay unindexed for weeks, or when log files show crawlers spending their time on parameters, filters and pagination instead of the pages that matter.
---

# Crawl budget audit

Search engines allocate a finite amount of crawling to any site. On a catalog or a site with faceted navigation, that budget is routinely spent on near-duplicate URLs while the pages you care about wait. This skill finds where the budget is going and reclaims it.

## What you need first

- Server access logs for at least 14 days, or Search Console crawl stats
- A full URL inventory (sitemap or crawl export)
- robots.txt

## Method

1. Pull crawl hits per URL pattern from the logs, grouping by path segment and query parameter rather than by individual URL.
2. Rank the patterns by share of total crawl hits. Anything above 5% that is not a page you want ranked is a leak.
3. For each leak, decide the correct instrument: robots.txt Disallow for URLs that must never be fetched, canonical for duplicates that must stay reachable, noindex for pages that must be crawled but not listed, and parameter rules only where the first three do not fit.
4. Check for crawl traps: infinite calendars, session IDs in URLs, sort orders that generate unbounded combinations.
5. Compare crawl frequency against last-modified dates. Pages that change often but are crawled rarely need better internal linking, not a directive.
6. Re-measure after 14 days and confirm the reclaimed share landed on the pages you intended.

## What this produces

A ranked table of crawl leaks with the specific directive to apply to each, plus a before/after share of crawl spent on revenue pages.

## Where this goes wrong

- Blocking a URL in robots.txt that already carries links - it keeps its equity but can no longer pass it on
- Using noindex on pages you also blocked from crawling, so the noindex is never seen
- Treating crawl budget as a problem on a 200-page site, where it almost never is

---

From the QuQi skill library - https://www.quqi.io/skills/crawl-budget-audit
Free to download · no account, no email

What you need first

  • Server access logs for at least 14 days, or Search Console crawl stats
  • A full URL inventory (sitemap or crawl export)
  • robots.txt

Method

  1. 01 Pull crawl hits per URL pattern from the logs, grouping by path segment and query parameter rather than by individual URL.
  2. 02 Rank the patterns by share of total crawl hits. Anything above 5% that is not a page you want ranked is a leak.
  3. 03 For each leak, decide the correct instrument: robots.txt Disallow for URLs that must never be fetched, canonical for duplicates that must stay reachable, noindex for pages that must be crawled but not listed, and parameter rules only where the first three do not fit.
  4. 04 Check for crawl traps: infinite calendars, session IDs in URLs, sort orders that generate unbounded combinations.
  5. 05 Compare crawl frequency against last-modified dates. Pages that change often but are crawled rarely need better internal linking, not a directive.
  6. 06 Re-measure after 14 days and confirm the reclaimed share landed on the pages you intended.

What this produces

A ranked table of crawl leaks with the specific directive to apply to each, plus a before/after share of crawl spent on revenue pages.

Where this goes wrong

  • Blocking a URL in robots.txt that already carries links - it keeps its equity but can no longer pass it on
  • Using noindex on pages you also blocked from crawling, so the noindex is never seen
  • Treating crawl budget as a problem on a 200-page site, where it almost never is

Use this skill in your own AI

The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.

Claude Code Save it as ~/.claude/skills/crawl-budget-audit/SKILL.md and Claude loads it on its own when what you are doing matches the trigger line. Put it in .claude/skills inside a project instead if the whole team should have it.
Claude Upload the file in the skills section of your settings. Once it is there it applies itself in any conversation where the trigger fits, so you do not have to remember it exists.
ChatGPT There is no skills format to install into, so paste the file contents into a Project instruction or a Custom GPT instead. It then applies to every chat in that project rather than only the one you paste it into.
Anything else Paste the markdown into the chat before your question. It works in any assistant, it just has to be pasted again each time.

Questions about this skill

When is a crawl budget audit the right method rather than just running a site crawler?

When a large site leaves pages unindexed for weeks and you need to know where the fetching actually went. A site crawler tells you what exists and what links to it; the logs tell you what Googlebot chose to spend its time on. On a few hundred URLs skip this entirely, because crawl budget is almost never the binding constraint at that size.

What do I need in hand before starting?

Fourteen days of server access logs or Search Console crawl stats, a full URL inventory from a sitemap or crawl export, and the current robots.txt. Without the inventory you cannot tell a leak from a page you meant to be crawled heavily. With three days of logs you read a burst as a pattern, because crawling arrives unevenly across a week.

What do I end up with, and which part gets used?

A table of URL patterns ranked by share of crawl hits, each with the directive to apply, plus a before and after share of crawl landing on revenue pages. The instrument column is the part that gets used: choosing between a robots.txt Disallow, a canonical and a noindex per pattern is the actual decision, and the ranking only says which to argue about first.

What is the mistake that ruins this, and what does it cost?

Blocking a pattern in robots.txt and putting a noindex on the same URLs. The crawler can no longer fetch the page, so the noindex is never read and the URLs sit in the index indefinitely. You have a fix that looks shipped and is not, and the same patterns reappear at the next audit. A blocked URL also stops passing on any equity it already carries.

More in Technical SEO