Robots.txt Validator
Lint Google-supported directives with source line numbers
Google-oriented robots.txt text lint with original line numbers and consecutive user-agent grouping. Only allow/disallow count as crawl rules. This does not simulate matching priority or determine live crawling/indexing. Input limit: 500 KiB and 20,000 lines.
This tool processes input in the page, does not request embedded addresses, and does not save input drafts. Site analytics records actions and summary information only.
The full guide also includes pitfalls, worked examples, snippets, FAQs, and related tools for checking results or troubleshooting.
About this tool
Review robots.txt text using Google-oriented lint rules. The report retains source line numbers, groups consecutive user-agent lines correctly, counts allow/disallow rules, and checks sitemap HTTP(S) URLs and duplicates. Missing separators are errors; unsupported Google directives such as noindex, crawl-delay, host and clean-param receive contextual warnings. Other crawlers may support some of those directives. Sitemap-only and comment-only files are allowed; empty allow/disallow values add no path restriction. The tool does not rewrite your file, simulate URL matching or check live fetching, HTTP status or indexing. Disallow does not guarantee removal from search. Input is limited to 500 KiB and 20,000 lines. Input drafts are not saved, and embedded URLs are not requested; site analytics records actions and counts.
Failure Clinic (Common Pitfalls)
Using Noindex in robots.txt to remove a page from search
Cause: Google does not support this directive in robots.txt. A disallowed URL can still be indexed.
Fix: Use robots meta or X-Robots-Tag on a page Google can crawl, then verify the result in the relevant search tools.
Treating crawl-delay as an error for every crawler
Cause: Google ignores crawl-delay; other crawlers may support it.
Fix: Interpret the warning in the context of the target crawler. This report is specifically Google-oriented.
Scenario Recipes
Review a robots file without losing line numbers
Goal: Check the Google-supported directive subset and inspect group boundaries.
- Paste the file with its original comments and blank lines.
- Read the reported source line and group for each directive; consecutive user-agent lines before a crawl rule share one group.
- Fix malformed lines and sitemap URLs, then review Google-unsupported directives separately from other crawlers’ behavior.
Result: A source-line report. The tool does not rewrite the file or simulate whether a particular URL is allowed.
Production Snippets
Two user agents can belong to one group
text
# Original line 1
User-agent: Googlebot
User-agent: Bingbot
Disallow: /private/
Sitemap: https://example.test/sitemap.xml
This is one group with one crawl rule.
The Sitemap directive does not terminate a group.Frequently Asked Questions
How are user-agent groups counted?
Consecutive user-agent lines, including intervening comments or blank lines, share a group until an allow/disallow rule has appeared. A later user-agent then begins another group. Sitemap and unknown directives do not terminate groups.
Can a file contain only Sitemap declarations?
Yes. Sitemap-only and comment-only files do not need invented user-agent groups. The report notes that no allow/disallow restrictions were found.
Can robots.txt noindex remove a Google result?
Google does not support noindex in robots.txt. Use robots meta or X-Robots-Tag where Google can fetch the page, and verify indexing separately.
Is crawl-delay invalid for every crawler?
No. Google ignores it, while some other crawlers support it. Warnings here identify the Google-specific boundary.
Does the report predict whether a URL will be crawled?
No. It does not simulate wildcard matching, rule precedence, crawler selection, redirects, caching or HTTP failures.
Does the tool upload or keep the file?
The checker processes the text locally, does not request its sitemap URLs and does not save input drafts. Site analytics records actions and counts.
Keep browsing