Should AI crawlers access your site?
Blocking AI crawlers in robots.txt sounds like a simple privacy decision. It is not. Some bots train models, some power search, and blocking the wrong one can remove you from AI answers. Here is what each does and how to set it up.
Many sites added Disallow: / for GPTBot in 2023 and never looked at it again. Some went further and blocked every AI user agent they could find.
That is a legitimate choice. But it is often made without knowing that the crawlers do different jobs. Block the training bot and you keep your content out of training. Block the search bot and you may vanish from AI answers.
This guide separates the two so you can decide on purpose.
The core idea: one company, several bots
OpenAI, Anthropic and Perplexity each run more than one user agent. Their own documentation splits them by purpose, and each is controlled separately in robots.txt.
OpenAI
Per OpenAI's crawler documentation, there are these robots:
- GPTBot crawls content that may be used to train OpenAI's generative AI foundation models. Disallowing it signals that your content should not be used for training.
- OAI-SearchBot surfaces websites in ChatGPT's search features. OpenAI says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links.
- ChatGPT-User visits pages for certain user actions in ChatGPT and Custom GPTs. OpenAI says it is not used to crawl automatically, and because the actions are user-initiated, robots.txt rules may not apply.
- OAI-AdsBot validates pages submitted as ads on ChatGPT. It is not used to train models.
OpenAI states that each setting is independent. You can allow OAI-SearchBot and disallow GPTBot. It also says robots.txt changes can take about 24 hours to take effect for search.
Anthropic
Anthropic's help center lists three:
- ClaudeBot collects web content that could contribute to model training. Restricting it signals that your future content should be excluded from training datasets.
- Claude-User fetches pages when a person asks Claude a question. Disabling it may reduce your visibility for user-directed web search.
- Claude-SearchBot navigates the web to improve search result quality. Disabling it prevents indexing for search optimization and may reduce your visibility in user search results.
Anthropic says its bots honor robots.txt and support the non-standard Crawl-delay directive. It also warns that blocking by IP address may not work reliably, because it stops Anthropic reading your robots.txt.
Perplexity
Perplexity's documentation lists two:
- PerplexityBot surfaces and links sites in Perplexity search results. Perplexity says it is not used to crawl content for AI foundation models.
- Perplexity-User supports user actions. Perplexity says it generally ignores robots.txt rules, since a user requested the fetch.
Training bots and search bots are a different decision
Here is the useful way to think about it.
Training crawlers (GPTBot, ClaudeBot) feed future models. Blocking them is about control over how your content is used. It has no documented effect on whether you appear in live search answers.
Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) feed the retrieval side. Blocking them takes you out of the pool of sources those products can surface and link.
User-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User) act on behalf of one person asking one question. Some vendors say robots.txt may not apply to them.
If your goal is to be recommended by AI assistants, allow the search crawlers. Whether to allow the training crawlers is a separate business call.
How to decide
Ask three questions.
1. Do we want AI assistants to recommend us? For most businesses selling something, yes. Allow the search bots.
2. Is our content the product? Publishers, paid research firms and licensed-content sites may reasonably block training bots. Their content is the thing they sell.
3. Are we comfortable with training use? For a software company with public marketing pages, many teams decide the upside of being known to models outweighs the cost. Others disagree. Both are defensible. Note that blocking training bots only affects future crawling, not data already collected.
There is no universal right answer. There is only an answer you chose knowingly.
Example robots.txt files
Each group starts with a User-agent line. Rules apply to the agent named in that group.
Allow everything AI
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Perplexity-User
Allow: /
If you have no rule blocking these agents, they are generally allowed by default. Listing them is optional but makes intent explicit.
Appear in AI search, opt out of training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
This matches each vendor's own description: training bots blocked, search bots allowed. Be aware that it is your reading of their documentation. Vendors can change their bots, so recheck now and then.
Block a private area only
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: PerplexityBot
Disallow: /account/
Disallow: /internal/
Stacking several User-agent lines above one set of rules is valid. Most sites should do this rather than block everything.
Common mistakes
A blanket User-agent: * block. This blocks every bot without a more specific group, including all the search crawlers above. Check any wildcard rules.
Blocking at the firewall. A CDN or WAF bot rule can return 403 to crawlers that robots.txt allows. Perplexity and OpenAI both publish IP ranges so you can allowlist them. Anthropic asks that you do not rely on IP blocking to opt out.
Typos in user-agent names. The names are GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot and Perplexity-User. A misspelled name matches nothing and fails silently.
Assuming robots.txt stops everything. It is a request, not a lock. Vendors say their bots honor it, and some user-triggered fetchers may not. If content must be private, put it behind a login.
Forgetting subdomains. Robots.txt applies per host. Anthropic notes you should do this for every subdomain.
Test what you set up
Do not trust a file you have not tested. The AI crawler access checker reads your robots.txt and shows which AI user agents are allowed or blocked. Run it after every change, and again after any CDN or firewall update.
Then check your server logs for the user-agent names above. If you allowed them and never see them, something between the internet and your server is in the way.
Also consider llms.txt
Robots.txt says what bots may fetch. llms.txt is a proposed plain-text file that points language models to your most useful pages. Treat it as a low-cost extra, not a guaranteed fix. If you want one, the llms.txt generator builds a starting file from your site.
Then measure the outcome
Changing robots.txt is an input. What you care about is the output: do AI assistants name and cite you? After you settle your rules, track it. Pineprompt runs your prompts across ChatGPT, Gemini, Perplexity, Claude, Google AI Overviews, AI Mode, Grok and Copilot, and reports where you appear and which sources are cited. Plans start at $99 a month with every engine included. See the pricing page.
Track your brand across every AI platform.
Pineprompt monitors eight AI platforms daily, in the country and language your buyers use. See our methodology for what we capture and how.