Sign in Start free
SEO

Robots and sitemap review

Use when you need someone to read a robots.txt and sitemap setup line by line for mistakes.

robots-sitemap-review.md
Download .md
You are reviewing crawl directives for errors and unintended consequences.

robots.txt contents: {{ROBOTS_TXT}}
Sitemap index and child sitemap structure: {{SITEMAP_STRUCTURE}}
Sample of URLs in the sitemaps: {{SITEMAP_SAMPLE}}
URL patterns that should NOT be indexed: {{EXCLUDE_PATTERNS}}
Site size (approximate URL count): {{SITE_SIZE}}

Output three sections.

A. robots.txt line review - a table: Line | What it does | Verdict (Correct / Risky / Broken) | Consequence if left | Suggested replacement.

B. Sitemap review - check for: URLs that are blocked in robots.txt, non-canonical URLs, non-200 URLs implied by the patterns, redirect chains, missing lastmod, files over the 50,000 URL or 50MB limit, and whether the structure allows useful indexation reporting.

C. Recommended robots.txt, rewritten in full, with a comment line above each rule explaining it.

Rules:
- Flag any Disallow that would also block a resource needed for rendering.
- State clearly where a Disallow has been used when noindex was the correct tool, and explain the difference in one sentence.
- Do not assume URLs exist that are not shown. If a check needs data you do not have, list it under "Cannot verify from what was supplied".

Fill in before running

Replace each placeholder with your own detail. The more specific you are, the less the model invents.

  • {{ROBOTS_TXT}}
  • {{SITEMAP_STRUCTURE}}
  • {{SITEMAP_SAMPLE}}
  • {{EXCLUDE_PATTERNS}}
  • {{SITE_SIZE}}

Getting a better result

  1. Paste the raw file, not a paraphrase - directive order and wildcards are the whole point.
  2. Include the exclude patterns explicitly, otherwise the review cannot tell intent from accident.
  3. Split sitemaps by template before reviewing so indexation reporting is actually diagnostic.

Questions about this prompt

When is a line-by-line review worth the time?

Before a migration, after a replatform, or whenever indexation numbers do not match what you expect. Robots and sitemap mistakes are cheap to make and expensive to notice, because the symptom is pages quietly not being indexed rather than an error anyone sees.

Why does it want the raw file?

Because directive order and wildcards are the entire point. A paraphrase of a robots.txt loses exactly the detail that causes the bug, and the most common serious mistakes are a wildcard matching more than intended or a rule ordered so it never applies.

Why include the exclude patterns explicitly?

So the review can tell intent from accident. A blocked directory is either deliberate or a mistake, and without knowing which you wanted, a reviewer can only guess and will flag both the same way.

Does splitting sitemaps actually help rankings?

Not directly, but it makes indexation reporting diagnostic instead of decorative. Splitting by template lets you see that product pages are indexed at ninety per cent and blog posts at thirty, which is a finding you cannot get from one combined file.