robots.txt generator
Build a robots.txt from multiple user agents, Allow / Disallow entries, Crawl-delay, Sitemap and Host — validated as you type, with a crawler cheat sheet and presets.
Runs in your browserEvery computation happens in your browser — your data never leaves this device.
Presets
Crawler rules
Sitemap and Host
Validation
Crawler cheat sheet
GooglebotGoogle SearchBingbotBing SearchBaiduspiderBaidu SearchSogou web spiderSogou Search360Spider360 SearchYandexBotYandex SearchDuckDuckBotDuckDuckGo SearchApplebotApple SearchPetalBotHuawei Petal SearchfacebookexternalhitFacebook crawlerTwitterbotX / Twitter cardsGPTBotOpenAI training crawler · AI training / crawlingChatGPT-UserChatGPT browsing · AI training / crawlingOAI-SearchBotOpenAI search index · AI training / crawlingClaudeBotAnthropic Claude crawler · AI training / crawlinganthropic-aiAnthropic training crawler · AI training / crawlingPerplexityBotPerplexity crawler · AI training / crawlingCCBotCommon Crawl dataset · AI training / crawlingGoogle-ExtendedGoogle AI training opt-out · AI training / crawlingApplebot-ExtendedApple AI training opt-out · AI training / crawlingBytespiderByteDance crawler · AI training / crawlingmeta-externalagentMeta AI crawler · AI training / crawlingAmazonbotAmazon crawler · AI training / crawlingMJ12botMajestic SEO crawler · AI training / crawlingAhrefsBotAhrefs SEO crawlerSemrushBotSemrush SEO crawlerrobots.txt
User-agent: * Allow: /
Statistics
1102What this tool does
- Ship a robots.txt before launch: allow the whole site but block admin paths, or block everything, using a preset as the starting point.
- Do not want your content scraped for AI training? The "block AI crawlers" preset writes a rule for GPTBot, ClaudeBot, CCBot, Bytespider and friends.
- Paths and URLs are validated as you type: paths must start with / and contain no spaces, and Sitemap and Host must be full http(s) URLs, with immediate feedback when they are not.
- Configure several rule sets for the same user agent and they merge with deduplication, so you never ship conflicting duplicate rules.
Example
Input
Rule group: User-agent *, Allow /, Disallow /admin/ and /api/; Sitemap https://example.com/sitemap.xml
Output
User-agent: * Allow: / Disallow: /admin/ Disallow: /api/ Sitemap: https://example.com/sitemap.xml
Directives inside a group are emitted in the order user agent, Allow, Disallow, Crawl-delay; Sitemap and Host form their own block after every group, separated by a blank line.
Frequently asked questions
Does Disallow really stop a path from being crawled?
No, not guaranteed. robots.txt is a voluntary convention that only well-behaved crawlers honour, and malicious ones ignore it entirely. It also cannot remove an already indexed page from search results, so protect sensitive content with authentication or server-level restrictions.
Which rule wins when Allow and Disallow conflict?
Per the spec the more specific (longer) matching path wins, and when lengths tie Allow takes precedence. That is why the usual pattern is to allow everything and then disallow admin directories, using Allow to carve out a subpath that was caught by mistake.
Which crawlers does the AI-blocking preset cover?
GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, anthropic-ai, PerplexityBot, CCBot, Google-Extended, Applebot-Extended, Bytespider, meta-externalagent, Amazonbot and MJ12bot each get a Disallow: / rule, plus a User-agent: * group that allows everyone else.
What happens if a path is invalid?
The validation panel lists the problem and the generated output skips the invalid entry. Paths must start with / and contain no spaces, Crawl-delay must be a non-negative number, and Sitemap and Host must be full http or https URLs.
Where does the generated file go?
In the site root, named exactly robots.txt, so it is reachable at https://your-domain/robots.txt. A robots.txt inside a subdirectory has no effect, and check that your CDN or reverse proxy does not block it.
Keywords:robots.txtrobotscrawleruser-agentsitemapseorobots.txt 生成爬虫禁止抓取AI 爬虫