Sign in Start free

Robots.txt Generator

Write crawl rules search engines actually follow, with your sitemap wired in.

Instant No signup

  

robots.txt controls crawling, not indexing. A blocked page can still appear in results if other sites link to it - use a noindex meta tag to keep a page out.

What it checks

  • Writes a User-agent: * group from one of three presets: allow everything, block common app paths, or block the whole site for staging.
  • The app preset disallows /admin/, /cart/, /checkout/, /account/, /search and sorted URLs matching /*?*sort=.
  • Adds a Disallow line for each extra path you list, and adds the leading slash if you left it off.
  • Optionally adds groups that block GPTBot, CCBot, ClaudeBot, Google-Extended, anthropic-ai and PerplexityBot from the whole site.
  • Adds a Sitemap line pointing at /sitemap.xml on the site you enter.
  • Builds the file in your browser, ready to copy or download as robots.txt.

Reading the result

Each Disallow value matches from the start of the path. /search blocks /search?q=shoes, and it also blocks /search-tips and /searchable-products. Read the list against your own URLs and make sure no page you want found starts with a blocked path.

The Sitemap line assumes your sitemap is at /sitemap.xml. If yours is elsewhere, such as /sitemap_index.xml, change that line before you upload. The file only works at the root of the host, for example https://example.com/robots.txt, and each subdomain needs its own.

The AI crawler option blocks the six bots named in the file and no others. Some of them collect training data and some fetch pages for assistants that cite their sources, so blocking all six can reduce how often you are quoted. Leave them allowed if being cited in AI answers matters to you.

Questions

Does this check my current robots.txt?

No. It writes a new file from the options you pick and does not fetch your site. Compare the result with your live file before you replace it, so you keep any rules that were added on purpose.

How do I keep a staging site out of search results?

Put it behind a password. The staging preset asks crawlers to stay out, but robots.txt is a request, and a blocked URL can still be listed if something links to it.

Do the AI bots also follow the rules under User-agent: *?

No. A crawler follows the most specific group that names it and ignores the others, so each AI bot group stands on its own. If you add a group for another bot by hand, repeat any Disallow lines you want it to follow.

Do I need a robots.txt at all?

No. Without one, crawlers assume they may fetch everything. It is still useful for pointing crawlers to your sitemap and keeping them out of carts, accounts and internal search results.

OR STOP DOING IT BY HAND

This checks one page, once. Your crew checks every page, every week, and fixes what it finds.

Add your site - free