Crawl budget audit
Search engines allocate a finite amount of crawling to any site. On a catalog or a site with faceted navigation, that budget is routinely spent on near-duplicate URLs while the pages you care about wait. This skill finds where the budget is going and reclaims it.
Use when a large site has pages that stay unindexed for weeks, or when log files show crawlers spending their time on parameters, filters and pagination instead of the pages that matter.
The skill file
What you need first
- Server access logs for at least 14 days, or Search Console crawl stats
- A full URL inventory (sitemap or crawl export)
- robots.txt
Method
- 01 Pull crawl hits per URL pattern from the logs, grouping by path segment and query parameter rather than by individual URL.
- 02 Rank the patterns by share of total crawl hits. Anything above 5% that is not a page you want ranked is a leak.
- 03 For each leak, decide the correct instrument: robots.txt Disallow for URLs that must never be fetched, canonical for duplicates that must stay reachable, noindex for pages that must be crawled but not listed, and parameter rules only where the first three do not fit.
- 04 Check for crawl traps: infinite calendars, session IDs in URLs, sort orders that generate unbounded combinations.
- 05 Compare crawl frequency against last-modified dates. Pages that change often but are crawled rarely need better internal linking, not a directive.
- 06 Re-measure after 14 days and confirm the reclaimed share landed on the pages you intended.
What this produces
A ranked table of crawl leaks with the specific directive to apply to each, plus a before/after share of crawl spent on revenue pages.
Where this goes wrong
- Blocking a URL in robots.txt that already carries links - it keeps its equity but can no longer pass it on
- Using noindex on pages you also blocked from crawling, so the noindex is never seen
- Treating crawl budget as a problem on a 200-page site, where it almost never is
Use this skill in your own AI
The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.
Questions about this skill
When is a crawl budget audit the right method rather than just running a site crawler?
When a large site leaves pages unindexed for weeks and you need to know where the fetching actually went. A site crawler tells you what exists and what links to it; the logs tell you what Googlebot chose to spend its time on. On a few hundred URLs skip this entirely, because crawl budget is almost never the binding constraint at that size.
What do I need in hand before starting?
Fourteen days of server access logs or Search Console crawl stats, a full URL inventory from a sitemap or crawl export, and the current robots.txt. Without the inventory you cannot tell a leak from a page you meant to be crawled heavily. With three days of logs you read a burst as a pattern, because crawling arrives unevenly across a week.
What do I end up with, and which part gets used?
A table of URL patterns ranked by share of crawl hits, each with the directive to apply, plus a before and after share of crawl landing on revenue pages. The instrument column is the part that gets used: choosing between a robots.txt Disallow, a canonical and a noindex per pattern is the actual decision, and the ranking only says which to argue about first.
What is the mistake that ruins this, and what does it cost?
Blocking a pattern in robots.txt and putting a noindex on the same URLs. The crawler can no longer fetch the page, so the noindex is never read and the URLs sit in the index indefinitely. You have a fix that looks shipped and is not, and the same patterns reappear at the next audit. A blocked URL also stops passing on any equity it already carries.
More in Technical SEO
Indexation gap analysis
Use when the number of pages you publish and the number Google indexes do not match, and you nee...
Core Web Vitals triage
Use when field data shows failing LCP, INP or CLS and you need to know which fix will actually m...
JavaScript rendering check
Use when a site renders content client-side and you need to confirm search engines actually see...
Redirect chain cleanup
Use after a migration, a domain change, or whenever a crawl reports redirects pointing at other...