# CCBot

## Quick answer

CCBot is the crawler run by Common Crawl, a nonprofit that publishes a free, public archive of the web. Many research teams and AI labs use that archive as training or research data, making CCBot a widely reused data source rather than a single company's assistant.

## What it does

CCBot crawls the public web to build Common Crawl's open dataset, a large, freely available archive that any researcher, company, or AI lab can download and use. Because the dataset is public, content CCBot gathers can end up as training or research data for many different organizations rather than a single one, which makes it different from a crawler run by one AI company for its own models. CCBot identifies itself in its user agent string as CCBot/2.0, and Common Crawl says it is aware that other crawlers sometimes falsely claim to be CCBot, so it recommends site owners verify the actual user agent string before assuming traffic is genuine. Disallowing CCBot has no effect on commercial search engines like Google or Bing, since those run their own separate crawlers.

## Facts

- Operator: Common Crawl
- Robots token: CCBot
- User agent string(s): CCBot/2.0 (https://commoncrawl.org/faq/)
- Purpose: training
- Respects robots.txt: yes
- IP ranges: https://index.commoncrawl.org/ccbot.json

## Should you block it?

Because Common Crawl's archive is open and reused by many different research groups and AI labs, blocking CCBot is less about any single product's visibility and more about whether a site wants its content available as free, public training and research data at all. A publisher protecting original reporting or content it monetizes directly might disallow CCBot specifically because that data flows to parties beyond Common Crawl itself, with no way to track who downloads it afterward. A brand hoping to be mentioned in AI answers should keep in mind that many AI models have learned from Common Crawl historically, so blocking CCBot going forward mainly affects future training snapshots rather than removing a brand from models already trained on earlier archives. Disallowing CCBot has no bearing on a site's presence in commercial search engines, since those use entirely separate crawlers.

## How to block it

Add this to robots.txt to block CCBot:

User-agent: CCBot
Disallow: /

Because other crawlers sometimes falsely identify themselves as CCBot, Common Crawl recommends verifying the user agent string of incoming traffic before relying on this rule alone.

## FAQ

### Who uses the Common Crawl archive that CCBot builds?

Many research teams and AI labs use Common Crawl's free, public dataset for training and research, rather than one company alone.

### Can other bots pretend to be CCBot?

Yes. Common Crawl says it is aware of crawlers falsely identifying themselves as CCBot and recommends verifying the user agent string.

### Does blocking CCBot affect Google or Bing search results?

No. CCBot feeds the separate Common Crawl archive and has no effect on commercial search engines, which crawl independently.

## Sources

- CCBot identifies itself in its UserAgent string as CCBot/2.0. Source: https://commoncrawl.org/ccbot
- Common Crawl is aware of other crawlers falsely identifying themselves as CCBot and recommends verifying the UserAgent string. Source: https://commoncrawl.org/ccbot
