Robots.txt checker
See a site's robots.txt rules for each crawler, spot mistakes, and test whether a page may be crawled.
Choosing a check opens its page. With the keyboard, use the arrow keys, then press Enter.
What we check
We fetch /robots.txt from the site's address (following up to the usual number of redirects) and explain it using the Robots Exclusion Protocol, RFC 9309:
- What the answer means: a 2xx answer means the rules apply; 4xx means crawlers may crawl everything; 5xx or no answer means they should crawl nothing until it is fixed.
- Groups: the rules for each user agent, and the
Sitemaplines. - Mistakes: blocking every crawler, rules outside a group, unknown lines, non-standard
Crawl-delay, relative sitemap addresses, blocking style sheet or script folders, and the wrong content type.
Test a path
Enter a path such as /blog/page and a user agent. We use the group for that crawler (or *), find the longest matching rule (an Allow wins a tie), and tell you which line decided. Search engines may differ in details beyond the standard.
robots.txt asks crawlers not to visit pages. It doesn't stop people opening them, and it isn't a way to keep a page out of search results.
Frequently asked questions
What does robots.txt do?
It tells well-behaved crawlers, such as search engines, which parts of a site they may crawl (RFC 9309). It is a request, not a security control: anyone can still visit the pages, and pages blocked from crawling can still appear in search results if other sites link to them.
How is the deciding rule chosen?
Crawlers use the group for their own name, or the "*" group if there isn't one. Within it, the rule with the longest matching path wins; if an Allow and a Disallow match equally, Allow wins. "*" matches any characters and "$" marks the end of the address.
What if there is no robots.txt?
If the server answers with a 4xx status such as 404, crawlers may treat the site as having no restrictions. If it answers with a 5xx error or can't be reached, crawlers should assume nothing may be crawled until it is fixed.
Does Crawl-delay work?
It isn't part of the standard. Some crawlers use it, but Google ignores it. Use your search engine's own tools to change crawl rate.
How much of the file is read?
Up to 512 KB. RFC 9309 asks crawlers to read at least 500 KiB, so rules after that may be ignored.
Related tools
-
Sitemap Checker
Find a site's XML sitemap, count its addresses and spot problems that make search engines ignore entries.
-
Website Down Checker
Check whether a website is responding right now, from our server: status, response time, redirects, IP addresses and HTTPS.
-
HTTP Header Checker
See the HTTP status and every response header a website sends, for the final page and each redirect on the way.
-
Redirect Checker
Follow a web address through every redirect, with the status code, destination and timing of each hop.
-
Website Health Check
Run the main website, DNS and email checks on one site at once and get a short report, area by area, with links to the full results.