Skip to content
MarketHQ
AI crawler

CCBot

Common Crawl's ai model training crawler.

Quick answer

CCBot is the crawler run by Common Crawl, a nonprofit that publishes a free, public archive of the web. Many research teams and AI labs use that archive as training or research data, making CCBot a widely reused data source rather than a single company's assistant.

What it does

CCBot crawls the public web to build Common Crawl's open dataset, a large, freely available archive that any researcher, company, or AI lab can download and use. Because the dataset is public, content CCBot gathers can end up as training or research data for many different organizations rather than a single one, which makes it different from a crawler run by one AI company for its own models. CCBot identifies itself in its user agent string as CCBot/2.0, and Common Crawl says it is aware that other crawlers sometimes falsely claim to be CCBot, so it recommends site owners verify the actual user agent string before assuming traffic is genuine. Disallowing CCBot has no effect on commercial search engines like Google or Bing, since those run their own separate crawlers.

Facts

Operator

Common Crawl

Robots token

CCBot

User agent string(s)

  • CCBot/2.0 (https://commoncrawl.org/faq/)

Purpose

AI model training

Respects robots.txt

Yes

IP ranges

https://index.commoncrawl.org/ccbot.json

From Common Crawl's documentation, checked Oct 2026.

Should you block it?

Because Common Crawl's archive is open and reused by many different research groups and AI labs, blocking CCBot is less about any single product's visibility and more about whether a site wants its content available as free, public training and research data at all. A publisher protecting original reporting or content it monetizes directly might disallow CCBot specifically because that data flows to parties beyond Common Crawl itself, with no way to track who downloads it afterward. A brand hoping to be mentioned in AI answers should keep in mind that many AI models have learned from Common Crawl historically, so blocking CCBot going forward mainly affects future training snapshots rather than removing a brand from models already trained on earlier archives. Disallowing CCBot has no bearing on a site's presence in commercial search engines, since those use entirely separate crawlers.

How to block it

Add this to robots.txt to block CCBot: User-agent: CCBot Disallow: / Because other crawlers sometimes falsely identify themselves as CCBot, Common Crawl recommends verifying the user agent string of incoming traffic before relying on this rule alone.

Frequently asked questions

Who uses the Common Crawl archive that CCBot builds?

Many research teams and AI labs use Common Crawl's free, public dataset for training and research, rather than one company alone.

Can other bots pretend to be CCBot?

Yes. Common Crawl says it is aware of crawlers falsely identifying themselves as CCBot and recommends verifying the user agent string.

Does blocking CCBot affect Google or Bing search results?

No. CCBot feeds the separate Common Crawl archive and has no effect on commercial search engines, which crawl independently.

Related

  • AI crawler directory — every AI crawler's robots.txt token, purpose and facts in one place.
  • Free AI crawler checker — check which AI crawlers a domain's robots.txt actually allows or blocks.
  • Glossary — definitions of the AI-visibility terms that come up alongside crawler behavior.

Sources

  • CCBot identifies itself in its UserAgent string as CCBot/2.0. Source: https://commoncrawl.org/ccbot (checked Oct 2026)
  • Common Crawl is aware of other crawlers falsely identifying themselves as CCBot and recommends verifying the UserAgent string. Source: https://commoncrawl.org/ccbot (checked Oct 2026)

MarketHQ tracks brand mentions across communities, news, blogs, social and AI answers, and turns them into gap analysis and action plans.