AI crawler list
18 crawlers and fetchers run by OpenAI, Anthropic, Google, Perplexity, Apple and others: what each one collects, the token to use in robots.txt, and whether it follows robots.txt. Every fact comes from the operator's own documentation, checked 2026-09-28.
| Crawler | Operator | Purpose | robots.txt token | Follows robots.txt |
|---|---|---|---|---|
| Amazonbot | Amazon | Mixed use | Amazonbot | Yes |
| Claude-SearchBot | Anthropic | AI search index | Claude-SearchBot | Yes |
| Claude-User | Anthropic | User-triggered fetch | Claude-User | Yes |
| ClaudeBot | Anthropic | AI training | ClaudeBot | Yes |
| Applebot-Extended | Apple | AI training | Applebot-Extended | Yes |
| Applebot | Apple | Mixed use | Applebot | Yes |
| Bytespider | ByteDance | AI search index | Bytespider | Not stated |
| CCBot | Common Crawl | AI training | CCBot | Yes |
| DuckAssistBot | DuckDuckGo | AI search index | DuckAssistBot | Yes |
| Google-Extended | Mixed use | Google-Extended | Yes | |
| GoogleOther | Mixed use | GoogleOther | Yes | |
| PetalBot | Huawei | AI search index | PetalBot | Yes |
| meta-externalagent | Meta | Mixed use | meta-externalagent | Yes |
| OAI-SearchBot | OpenAI | AI search index | OAI-SearchBot | Yes |
| ChatGPT-User | OpenAI | User-triggered fetch | ChatGPT-User | Not always |
| GPTBot | OpenAI | AI training | GPTBot | Yes |
| PerplexityBot | Perplexity | AI search index | PerplexityBot | Yes |
| Perplexity-User | Perplexity | User-triggered fetch | Perplexity-User | Not always |
What each purpose means
AI search index: Builds the index an AI search product answers and cites from. Block it and you drop out of those answers.
User-triggered fetch: Fetches a page live because someone asked the assistant about it. Operators say some of these may not follow robots.txt, since a person made the request.
AI training: Collects pages that may be used to train AI models. Blocking it does not, by itself, remove you from AI search answers.
Mixed use: Used for more than one of the above.
Block AI training, stay in AI search
These lines opt out of the 4 training crawlers above and leave search and user-triggered crawlers alone, so AI search products can still find and cite you. Add them to the robots.txt at the root of your site.
User-agent: GPTBot User-agent: ClaudeBot User-agent: Applebot-Extended User-agent: CCBot Disallow: /
Block every AI crawler
This removes you from AI search answers too. Crawlers marked "Not always" above may still fetch pages when a user asks for them.
User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-User User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Google-Extended User-agent: GoogleOther User-agent: Applebot User-agent: Applebot-Extended User-agent: Amazonbot User-agent: CCBot User-agent: meta-externalagent User-agent: Bytespider User-agent: PetalBot User-agent: DuckAssistBot Disallow: /
Are AI engines citing you?
Being crawlable is step one. Citedify checks whether ChatGPT, Claude, Perplexity and Google AI Overviews recommend your brand, and tells you what to fix. Next, read how to rank in Google AI Overviews.