Indexation gap analysis
Publishing is not indexing. This skill reconciles what exists, what is submitted, and what is actually in the index, then classifies every gap by cause so each one has an owner and a fix.
Use when the number of pages you publish and the number Google indexes do not match, and you need to know which specific pages are missing and why.
The skill file
What you need first
- Sitemap URLs
- Search Console Page Indexing export
- A crawl of the live site
Method
- 01 Build three sets: URLs on the site, URLs in the sitemap, URLs reported as indexed.
- 02 Compute the differences in both directions - indexed-but-not-in-sitemap matters as much as the reverse.
- 03 For each unindexed URL, read the Search Console reason verbatim rather than assuming: discovered-not-crawled, crawled-not-indexed, duplicate, soft 404 and noindex are five different problems with five different fixes.
- 04 Group by reason and count. The largest group is where the work is.
- 05 For crawled-not-indexed - usually the largest group - assess quality honestly: thin, duplicate or templated pages are a content decision, not a technical one.
- 06 Fix the technical causes first because they are unambiguous, then re-submit and wait a full crawl cycle before judging the content ones.
What this produces
Every unindexed URL classified by cause, with counts per cause and the specific remedy for each group.
Where this goes wrong
- Re-submitting a sitemap repeatedly instead of fixing the reason
- Assuming crawled-not-indexed is a bug - it is usually a quality judgment
- Counting site: results as an index count, which they are not
Use this skill in your own AI
The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.
Questions about this skill
When should I reconcile the three sets rather than just resubmitting the sitemap?
When your published count and your indexed count diverge and nobody can name which pages are missing. Resubmitting a sitemap or running a site: query gives you neither a list nor a cause, and site: results are an estimate rather than an index count. This produces a named reason per URL, which is what turns the gap into work someone can own.
What do I need in hand before starting?
Sitemap URLs, the Search Console Page Indexing export and a fresh crawl of the live site. Drop the crawl and you lose the indexed-but-not-in-sitemap direction, which is where stray parameter and filter URLs usually hide. Work from summary counts rather than the export and you lose the per-URL reason, at which point you are guessing between five problems that share one symptom.
What do I end up with, and which part gets used?
Every unindexed URL classified by the reason Search Console reports, with counts per reason and a remedy per group. The counts are what gets used, because the largest group is where the work is and the rest is a distraction. The second useful line is the split between technical causes, which are unambiguous, and crawled-not-indexed, which is a content decision.
What is the mistake that ruins this, and what does it cost?
Reading crawled-not-indexed as a bug and resubmitting the sitemap at it. In most cases the page was fetched, assessed and declined, and resubmission changes nothing about that judgment. The cost is a quarter spent on submission mechanics while the thin or templated pages causing it stay exactly as they are. Fix the technical causes first, then wait a full crawl cycle before judging the rest.
More in Technical SEO
Crawl budget audit
Use when a large site has pages that stay unindexed for weeks, or when log files show crawlers s...
Core Web Vitals triage
Use when field data shows failing LCP, INP or CLS and you need to know which fix will actually m...
JavaScript rendering check
Use when a site renders content client-side and you need to confirm search engines actually see...
Redirect chain cleanup
Use after a migration, a domain change, or whenever a crawl reports redirects pointing at other...