Skip to main content
WebTools

Robots.txt checker

See a site's robots.txt rules for each crawler, spot mistakes, and test whether a page may be crawled.

Enter a path such as /blog/ to see whether the user agent may crawl it.

The crawler to test: * for any crawler, or a name such as Googlebot.

Choosing a check opens its page. With the keyboard, use the arrow keys, then press Enter.

What we check

We fetch /robots.txt from the site's address (following up to the usual number of redirects) and explain it using the Robots Exclusion Protocol, RFC 9309:

Test a path

Enter a path such as /blog/page and a user agent. We use the group for that crawler (or *), find the longest matching rule (an Allow wins a tie), and tell you which line decided. Search engines may differ in details beyond the standard.

robots.txt asks crawlers not to visit pages. It doesn't stop people opening them, and it isn't a way to keep a page out of search results.

Frequently asked questions

What does robots.txt do?

It tells well-behaved crawlers, such as search engines, which parts of a site they may crawl (RFC 9309). It is a request, not a security control: anyone can still visit the pages, and pages blocked from crawling can still appear in search results if other sites link to them.

How is the deciding rule chosen?

Crawlers use the group for their own name, or the "*" group if there isn't one. Within it, the rule with the longest matching path wins; if an Allow and a Disallow match equally, Allow wins. "*" matches any characters and "$" marks the end of the address.

What if there is no robots.txt?

If the server answers with a 4xx status such as 404, crawlers may treat the site as having no restrictions. If it answers with a 5xx error or can't be reached, crawlers should assume nothing may be crawled until it is fixed.

Does Crawl-delay work?

It isn't part of the standard. Some crawlers use it, but Google ignores it. Use your search engine's own tools to change crawl rate.

How much of the file is read?

Up to 512 KB. RFC 9309 asks crawlers to read at least 500 KiB, so rules after that may be ignored.

  • Sitemap Checker

    Find a site's XML sitemap, count its addresses and spot problems that make search engines ignore entries.

  • Website Down Checker

    Check whether a website is responding right now, from our server: status, response time, redirects, IP addresses and HTTPS.

  • HTTP Header Checker

    See the HTTP status and every response header a website sends, for the final page and each redirect on the way.

  • Redirect Checker

    Follow a web address through every redirect, with the status code, destination and timing of each hop.

  • Website Health Check

    Run the main website, DNS and email checks on one site at once and get a short report, area by area, with links to the full results.