Sign in Start free
SEO

Log file crawl budget analysis

Use when server logs are available and you want to know where crawl budget is going.

log-file-crawl-analysis.md
Download .md
You are analyzing server log data for crawl behavior.

Site: {{SITE_URL}}
Log summary (URL or directory, bot hits, status codes, date range): {{LOG_SUMMARY}}
Site structure and which sections make money: {{SITE_STRUCTURE}}
Total indexable URL count: {{INDEXABLE_COUNT}}

Output:
1. A crawl distribution table: Section or pattern | Bot hits | Share of total crawl | Share of indexable URLs | Commercial value (High / Medium / Low) | Verdict (Under-crawled / Balanced / Wasted).
2. "Crawl waste" - the patterns absorbing crawl with no value, ordered by hits, each with the mechanism (parameters, faceted navigation, pagination, redirect chains, soft 404s, calendars, session ids).
3. "Never crawled but important" - value sections with low or zero bot hits, with the likely reason.
4. A fix list: Action | Method (robots, noindex, canonical, parameter handling, internal linking, removal) | Expected effect on crawl distribution.

Constraints:
- Do not confuse low crawl with low indexation. Say which the data actually shows.
- Verify bot identity is claimed in the data before treating hits as genuine search crawler traffic; if verification status is unknown, say so.
- Do not recommend blocking a pattern in robots.txt if the pages need to be de-indexed first. Explain the ordering.

Fill in before running

Replace each placeholder with your own detail. The more specific you are, the less the model invents.

  • {{SITE_URL}}
  • {{LOG_SUMMARY}}
  • {{SITE_STRUCTURE}}
  • {{INDEXABLE_COUNT}}

Getting a better result

  1. Aggregate by directory or pattern before pasting - raw log lines waste the context window.
  2. Use at least 30 days of logs so weekly crawl cycles do not distort the picture.
  3. Pair this with an indexation report to see which under-crawled sections are actually missing.

Questions about this prompt

When is log analysis worth doing?

On large sites where indexation is incomplete and you cannot see why. On a small site crawl budget is rarely the constraint, and the effort is better spent elsewhere. The signal you are looking for is Google spending its time on pages that do not matter.

How much log data do I need?

At least thirty days, aggregated by directory or pattern before you paste it. Shorter windows let weekly crawl cycles distort the picture, and raw log lines waste the context window without adding anything the aggregate does not show.

What does a bad result look like?

Crawl concentrated on parameter URLs, pagination or filtered pages while your commercial templates are visited rarely. That is budget being spent on pages you never wanted indexed, and it is usually fixable with the faceted navigation rules rather than with more content.

What should I pair it with?

An indexation report. Logs tell you what was crawled and indexation tells you what made it in. Under-crawled sections that are also missing from the index are the priority; under-crawled sections that are indexed fine are not a problem.