What is CCBot?

Common Crawl's crawler, which builds a free, open repository of web crawl data that anyone can access and analyze.

Operator
Common Crawl
Purpose
AI training
Feeds
Common Crawl open web-crawl dataset
robots.txt token
CCBot
Follows robots.txt
Yes

This classification is an inference. Common Crawl's own pages describe an open repository of crawl data for 'research and analysis' and do not mention AI training.

CCBot user agent

CCBot/2.0 (https://commoncrawl.org/faq/)

Common Crawl says the version number may be incremented in the future and warns that other crawlers falsely identify themselves as CCBot.

Match it in robots.txt with the token CCBot. User agents can be faked, so check requests against Common Crawl's published IP ranges before trusting them. Reverse DNS on IPv4 should resolve to *.crawl.commoncrawl.org (e.g. 18-97-14-84.crawl.commoncrawl.org) and forward-resolve to the same IP; IPv6 has no reverse DNS yet.

Should you block CCBot?

A CCBot Disallow rule stops Common Crawl from crawling your site, and Common Crawl periodically rechecks robots.txt; publishers can also ask to join its opt-out registry.

Yes. Common Crawl says CCBot follows robots.txt. “Add these lines to your robots.txt file and our crawler will stop crawling your website”; Crawl-delay is honored.

Block CCBot

User-agent: CCBot
Disallow: /

Allow CCBot

User-agent: CCBot
Allow: /

CCBot FAQ

What is CCBot?

Common Crawl's crawler, which builds a free, open repository of web crawl data that anyone can access and analyze.

What is the CCBot user agent?

CCBot identifies itself as: CCBot/2.0 (https://commoncrawl.org/faq/). In robots.txt, match it with the token "CCBot".

Does CCBot respect robots.txt?

Yes. Common Crawl says CCBot follows robots.txt.

How do I block CCBot?

Add "User-agent: CCBot" followed by "Disallow: /" to the robots.txt file at the root of your site. A CCBot Disallow rule stops Common Crawl from crawling your site, and Common Crawl periodically rechecks robots.txt; publishers can also ask to join its opt-out registry.

Crawlable isn't the same as cited

Letting AI crawlers in is step one. Citedify checks whether ChatGPT, Claude, Perplexity and Google AI Overviews actually recommend your brand, and what to fix when they don't.

Facts on this page come from Common Crawl's documentation, checked 2026-09-28. Sources: commoncrawl.org/ccbot, commoncrawl.org/faq, commoncrawl.org/