Skip to content
UniKit

robots.txt generator

Build a robots.txt from multiple user agents, Allow / Disallow entries, Crawl-delay, Sitemap and Host — validated as you type, with a crawler cheat sheet and presets.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Presets

Crawler rules

Allow
/
Disallow

Sitemap and Host

Validation

Crawler cheat sheet

GooglebotGoogle Search
BingbotBing Search
BaiduspiderBaidu Search
Sogou web spiderSogou Search
360Spider360 Search
YandexBotYandex Search
DuckDuckBotDuckDuckGo Search
ApplebotApple Search
PetalBotHuawei Petal Search
facebookexternalhitFacebook crawler
TwitterbotX / Twitter cards
GPTBotOpenAI training crawler · AI training / crawling
ChatGPT-UserChatGPT browsing · AI training / crawling
OAI-SearchBotOpenAI search index · AI training / crawling
ClaudeBotAnthropic Claude crawler · AI training / crawling
anthropic-aiAnthropic training crawler · AI training / crawling
PerplexityBotPerplexity crawler · AI training / crawling
CCBotCommon Crawl dataset · AI training / crawling
Google-ExtendedGoogle AI training opt-out · AI training / crawling
Applebot-ExtendedApple AI training opt-out · AI training / crawling
BytespiderByteDance crawler · AI training / crawling
meta-externalagentMeta AI crawler · AI training / crawling
AmazonbotAmazon crawler · AI training / crawling
MJ12botMajestic SEO crawler · AI training / crawling
AhrefsBotAhrefs SEO crawler
SemrushBotSemrush SEO crawler

robots.txt

User-agent: *
Allow: /

Statistics

Rule groups1
Rule lines1
Sitemaps0
Lines2

What this tool does

  • Ship a robots.txt before launch: allow the whole site but block admin paths, or block everything, using a preset as the starting point.
  • Do not want your content scraped for AI training? The "block AI crawlers" preset writes a rule for GPTBot, ClaudeBot, CCBot, Bytespider and friends.
  • Paths and URLs are validated as you type: paths must start with / and contain no spaces, and Sitemap and Host must be full http(s) URLs, with immediate feedback when they are not.
  • Configure several rule sets for the same user agent and they merge with deduplication, so you never ship conflicting duplicate rules.

Example

Input

Rule group: User-agent *, Allow /, Disallow /admin/ and /api/; Sitemap https://example.com/sitemap.xml

Output

User-agent: *
Allow: /
Disallow: /admin/
Disallow: /api/

Sitemap: https://example.com/sitemap.xml

Directives inside a group are emitted in the order user agent, Allow, Disallow, Crawl-delay; Sitemap and Host form their own block after every group, separated by a blank line.

Frequently asked questions

Does Disallow really stop a path from being crawled?

No, not guaranteed. robots.txt is a voluntary convention that only well-behaved crawlers honour, and malicious ones ignore it entirely. It also cannot remove an already indexed page from search results, so protect sensitive content with authentication or server-level restrictions.

Which rule wins when Allow and Disallow conflict?

Per the spec the more specific (longer) matching path wins, and when lengths tie Allow takes precedence. That is why the usual pattern is to allow everything and then disallow admin directories, using Allow to carve out a subpath that was caught by mistake.

Which crawlers does the AI-blocking preset cover?

GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, anthropic-ai, PerplexityBot, CCBot, Google-Extended, Applebot-Extended, Bytespider, meta-externalagent, Amazonbot and MJ12bot each get a Disallow: / rule, plus a User-agent: * group that allows everyone else.

What happens if a path is invalid?

The validation panel lists the problem and the generated output skips the invalid entry. Paths must start with / and contain no spaces, Crawl-delay must be a non-negative number, and Sitemap and Host must be full http or https URLs.

Where does the generated file go?

In the site root, named exactly robots.txt, so it is reachable at https://your-domain/robots.txt. A robots.txt inside a subdirectory has no effect, and check that your CDN or reverse proxy does not block it.

Keywords:robots.txtrobotscrawleruser-agentsitemapseorobots.txt 生成爬虫禁止抓取AI 爬虫

Related tools