robots.txt for AI Crawlers: What to Allow, What to Block
How to configure robots.txt for GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot and Google-Extended, with copy-paste rules for every AI crawler policy.

Your robots.txt file has quietly become one of the most important files for AI search. A single line can decide whether ChatGPT, Perplexity and Claude can cite your site at all. Yet most robots.txt files were written years ago for Google, and many now block AI crawlers by accident, or block the wrong ones on purpose.
This guide explains which AI crawlers exist, what each one does, and exactly what to put in your robots.txt for the policy you want. It is part of our technical AEO guide, which covers the rest of the machine-readability checklist.
The key idea: search crawlers vs training crawlers
AI companies now run separate crawlers for separate jobs, and robots.txt controls each one independently. The distinction you need to understand is simple:
- Search crawlers fetch pages so the engine can show and cite them in answers. Block these and your site disappears from that engine's answers.
- Training crawlers collect content to train future models. Block these and your content stays out of training, but your visibility in AI search answers is unaffected.
- User agents fetch a page in real time because someone asked about it. These are not automated crawls, and robots.txt may not apply to them.
That means you do not have to choose between "visible in AI search" and "my content is not used for training". You can have both.
The AI crawlers that matter
| Token | Company | Type | What it controls |
|---|---|---|---|
OAI-SearchBot | OpenAI | Search | Inclusion in ChatGPT search answers |
GPTBot | OpenAI | Training | Use of your content to train OpenAI models |
ChatGPT-User | OpenAI | User agent | Pages fetched when a user asks ChatGPT |
PerplexityBot | Perplexity | Search | Surfacing and citing your site in Perplexity |
Perplexity-User | Perplexity | User agent | Pages fetched for a user's question |
Claude-SearchBot | Anthropic | Search | Claude's search results |
ClaudeBot | Anthropic | Training | Use of your content to train Claude |
Claude-User | Anthropic | User agent | Pages fetched when a user asks Claude |
Google-Extended | Training and grounding | Gemini training and grounding in Gemini apps | |
Applebot-Extended | Apple | Training | Use of content for Apple's AI models |
Two Google details are worth repeating because they cause so much confusion. Google-Extended is not a crawler that visits your site; it is a token Google checks when deciding how content its normal crawlers fetched may be used. And AI Overviews in Google Search come from regular Googlebot, so blocking Google-Extended does not remove you from AI Overviews or affect your rankings.
Copy-paste robots.txt configurations
Policy 1: Maximum AI visibility (recommended for most businesses)
If you want to be found and cited everywhere, allow everything and say so explicitly:
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Allow: /
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml
Policy 2: Visible in AI search, opted out of training
This is the most common choice for publishers and companies with proprietary content. You stay citable in ChatGPT, Perplexity and Claude answers while keeping content out of model training:
# AI search: allowed
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Allow: /
# AI training: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
User-agent: *
Allow: /
Policy 3: Block all AI crawlers
Only choose this if you have decided AI search is not a channel for you. Your site will not appear in ChatGPT search, Perplexity or Claude answers:
User-agent: OAI-SearchBot
User-agent: GPTBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
Common mistakes that block AI crawlers by accident
Grouped user agents. In robots.txt, consecutive User-agent lines share the rules that follow them. Teams often add a new bot to an existing "block scrapers" group without realizing that every agent in that group, including an AI search crawler someone added months ago, now shares the Disallow: /.
A site-wide wildcard block. User-agent: * followed by Disallow: / blocks every crawler that does not have its own group. This is common on staging sites and occasionally ships to production. Any AI crawler without an explicit group inherits the block.
Blocking GPTBot to "stay out of ChatGPT". As covered above, GPTBot is the training crawler. If the goal is to stay out of ChatGPT search, the token is OAI-SearchBot. If the goal is to appear in ChatGPT while staying out of training, blocking only GPTBot is exactly right.
Firewall rules that override robots.txt. A CDN or bot-protection rule that challenges or blocks unknown bots can stop AI crawlers even when robots.txt allows them. If you allow AI crawlers in robots.txt, check that your firewall allows them too.
How to check your current setup
Open https://yourdomain.com/robots.txt and read it top to bottom, grouping each block of User-agent lines with the rules beneath it. For each AI search crawler, ask: does it have its own group? If not, what does the * group say?
Or let a tool do it. The free AI crawler checker reads your robots.txt the way crawlers do and shows which AI crawlers are allowed or blocked, split into search, training and user agents. For the whole page, run a free AEOCheck scan and the AI crawler access check parses your robots.txt the way crawlers do, including grouped user agents and wildcard fallbacks, and tells you which AI search crawlers are blocked versus which are only opted out of training. The same scan covers llms.txt, schema, metadata and the rest of its 25 AEO and GEO checks. See an example in the sample report, or compare plans to monitor your pages over time.
What other sites do
We checked the robots.txt of 85 well-known sites. Of the 72 files we could read, 7 blocked at least one AI search crawler, mostly publishers and large retailers that block on purpose, and 13 blocked GPTBot. See the full numbers and method in the AI readiness study. Unless you have a licensing reason to opt out, a business that wants to be recommended should not block the search crawlers.