Sign in Start free

Indexation gap analysis

Publishing is not indexing. This skill reconciles what exists, what is submitted, and what is actually in the index, then classifies every gap by cause so each one has an owner and a fix.

CATEGORY
Technical SEO
FORMAT
indexation-gap-analysis.md
STEPS
6
PRICE
Free - no account
WHEN TO REACH FOR THIS

Use when the number of pages you publish and the number Google indexes do not match, and you need to know which specific pages are missing and why.

The skill file

indexation-gap-analysis.md
---
name: indexation-gap-analysis
description: Use when the number of pages you publish and the number Google indexes do not match, and you need to know which specific pages are missing and why.
---

# Indexation gap analysis

Publishing is not indexing. This skill reconciles what exists, what is submitted, and what is actually in the index, then classifies every gap by cause so each one has an owner and a fix.

## What you need first

- Sitemap URLs
- Search Console Page Indexing export
- A crawl of the live site

## Method

1. Build three sets: URLs on the site, URLs in the sitemap, URLs reported as indexed.
2. Compute the differences in both directions - indexed-but-not-in-sitemap matters as much as the reverse.
3. For each unindexed URL, read the Search Console reason verbatim rather than assuming: discovered-not-crawled, crawled-not-indexed, duplicate, soft 404 and noindex are five different problems with five different fixes.
4. Group by reason and count. The largest group is where the work is.
5. For crawled-not-indexed - usually the largest group - assess quality honestly: thin, duplicate or templated pages are a content decision, not a technical one.
6. Fix the technical causes first because they are unambiguous, then re-submit and wait a full crawl cycle before judging the content ones.

## What this produces

Every unindexed URL classified by cause, with counts per cause and the specific remedy for each group.

## Where this goes wrong

- Re-submitting a sitemap repeatedly instead of fixing the reason
- Assuming crawled-not-indexed is a bug - it is usually a quality judgment
- Counting site: results as an index count, which they are not

---

From the QuQi skill library - https://www.quqi.io/skills/indexation-gap-analysis
Free to download · no account, no email

What you need first

  • Sitemap URLs
  • Search Console Page Indexing export
  • A crawl of the live site

Method

  1. 01 Build three sets: URLs on the site, URLs in the sitemap, URLs reported as indexed.
  2. 02 Compute the differences in both directions - indexed-but-not-in-sitemap matters as much as the reverse.
  3. 03 For each unindexed URL, read the Search Console reason verbatim rather than assuming: discovered-not-crawled, crawled-not-indexed, duplicate, soft 404 and noindex are five different problems with five different fixes.
  4. 04 Group by reason and count. The largest group is where the work is.
  5. 05 For crawled-not-indexed - usually the largest group - assess quality honestly: thin, duplicate or templated pages are a content decision, not a technical one.
  6. 06 Fix the technical causes first because they are unambiguous, then re-submit and wait a full crawl cycle before judging the content ones.

What this produces

Every unindexed URL classified by cause, with counts per cause and the specific remedy for each group.

Where this goes wrong

  • Re-submitting a sitemap repeatedly instead of fixing the reason
  • Assuming crawled-not-indexed is a bug - it is usually a quality judgment
  • Counting site: results as an index count, which they are not

Use this skill in your own AI

The download is a plain markdown file with the name and trigger in its frontmatter. Where an assistant supports skills it can load itself, that frontmatter is what it reads to decide this one applies.

Claude Code Save it as ~/.claude/skills/indexation-gap-analysis/SKILL.md and Claude loads it on its own when what you are doing matches the trigger line. Put it in .claude/skills inside a project instead if the whole team should have it.
Claude Upload the file in the skills section of your settings. Once it is there it applies itself in any conversation where the trigger fits, so you do not have to remember it exists.
ChatGPT There is no skills format to install into, so paste the file contents into a Project instruction or a Custom GPT instead. It then applies to every chat in that project rather than only the one you paste it into.
Anything else Paste the markdown into the chat before your question. It works in any assistant, it just has to be pasted again each time.

Questions about this skill

When should I reconcile the three sets rather than just resubmitting the sitemap?

When your published count and your indexed count diverge and nobody can name which pages are missing. Resubmitting a sitemap or running a site: query gives you neither a list nor a cause, and site: results are an estimate rather than an index count. This produces a named reason per URL, which is what turns the gap into work someone can own.

What do I need in hand before starting?

Sitemap URLs, the Search Console Page Indexing export and a fresh crawl of the live site. Drop the crawl and you lose the indexed-but-not-in-sitemap direction, which is where stray parameter and filter URLs usually hide. Work from summary counts rather than the export and you lose the per-URL reason, at which point you are guessing between five problems that share one symptom.

What do I end up with, and which part gets used?

Every unindexed URL classified by the reason Search Console reports, with counts per reason and a remedy per group. The counts are what gets used, because the largest group is where the work is and the rest is a distraction. The second useful line is the split between technical causes, which are unambiguous, and crawled-not-indexed, which is a content decision.

What is the mistake that ruins this, and what does it cost?

Reading crawled-not-indexed as a bug and resubmitting the sitemap at it. In most cases the page was fetched, assessed and declined, and resubmission changes nothing about that judgment. The cost is a quarter spent on submission mechanics while the thin or templated pages causing it stay exactly as they are. Fix the technical causes first, then wait a full crawl cycle before judging the rest.

More in Technical SEO