AI crawler list

18 crawlers and fetchers run by OpenAI, Anthropic, Google, Perplexity, Apple and others: what each one collects, the token to use in robots.txt, and whether it follows robots.txt. Every fact comes from the operator's own documentation, checked 2026-09-28.

CrawlerOperatorPurposerobots.txt tokenFollows robots.txt
AmazonbotAmazonMixed useAmazonbotYes
Claude-SearchBotAnthropicAI search indexClaude-SearchBotYes
Claude-UserAnthropicUser-triggered fetchClaude-UserYes
ClaudeBotAnthropicAI trainingClaudeBotYes
Applebot-ExtendedAppleAI trainingApplebot-ExtendedYes
ApplebotAppleMixed useApplebotYes
BytespiderByteDanceAI search indexBytespiderNot stated
CCBotCommon CrawlAI trainingCCBotYes
DuckAssistBotDuckDuckGoAI search indexDuckAssistBotYes
Google-ExtendedGoogleMixed useGoogle-ExtendedYes
GoogleOtherGoogleMixed useGoogleOtherYes
PetalBotHuaweiAI search indexPetalBotYes
meta-externalagentMetaMixed usemeta-externalagentYes
OAI-SearchBotOpenAIAI search indexOAI-SearchBotYes
ChatGPT-UserOpenAIUser-triggered fetchChatGPT-UserNot always
GPTBotOpenAIAI trainingGPTBotYes
PerplexityBotPerplexityAI search indexPerplexityBotYes
Perplexity-UserPerplexityUser-triggered fetchPerplexity-UserNot always

What each purpose means

AI search index: Builds the index an AI search product answers and cites from. Block it and you drop out of those answers.

User-triggered fetch: Fetches a page live because someone asked the assistant about it. Operators say some of these may not follow robots.txt, since a person made the request.

AI training: Collects pages that may be used to train AI models. Blocking it does not, by itself, remove you from AI search answers.

Mixed use: Used for more than one of the above.

Block AI training, stay in AI search

These lines opt out of the 4 training crawlers above and leave search and user-triggered crawlers alone, so AI search products can still find and cite you. Add them to the robots.txt at the root of your site.

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Applebot-Extended
User-agent: CCBot
Disallow: /

Block every AI crawler

This removes you from AI search answers too. Crawlers marked "Not always" above may still fetch pages when a user asks for them.

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-User
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: GoogleOther
User-agent: Applebot
User-agent: Applebot-Extended
User-agent: Amazonbot
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Bytespider
User-agent: PetalBot
User-agent: DuckAssistBot
Disallow: /

Are AI engines citing you?

Being crawlable is step one. Citedify checks whether ChatGPT, Claude, Perplexity and Google AI Overviews recommend your brand, and tells you what to fix. Next, read how to rank in Google AI Overviews.