# ClaudeBot

## Quick answer

ClaudeBot is Anthropic's web crawler, which collects publicly available web content that could help train Anthropic's generative AI models, the kind that power Claude. It honors standard robots.txt directives plus a non-standard crawl-delay setting, and is separate from Claude-User and Claude-SearchBot.

## What it does

ClaudeBot crawls the web to collect content that could potentially contribute to training Anthropic's generative AI models, with the goal of improving their usefulness and safety. It operates separately from Claude-User, which fetches pages a live user or Claude's own browsing asked for, and from Claude-SearchBot, which supports search-style answers; each is controlled independently through its own robots.txt token. Anthropic's bots honor standard "do not crawl" signals in robots.txt, and Anthropic also supports the non-standard Crawl-delay directive, which lets a site slow down how fast ClaudeBot crawls rather than blocking it outright. A site that disallows ClaudeBot is opting its content out of future training data collection, without necessarily affecting whether Claude can browse or cite that page when a user explicitly asks about it elsewhere.

## Facts

- Operator: Anthropic
- Robots token: ClaudeBot
- User agent string(s): Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)
- Purpose: training
- Respects robots.txt: yes
- IP ranges: https://claude.com/crawling/bots.json

## Should you block it?

Disallowing ClaudeBot keeps a site's future content out of Anthropic's training datasets, which matters to publishers protecting original work, subscription content, or material they don't want reused to train a model. But Claude-User and Claude-SearchBot are separate from ClaudeBot, so blocking training crawling doesn't automatically stop Claude from fetching or citing a page when answering a specific user question, nor does it guarantee removal from how Claude discusses a brand based on content it already learned elsewhere. A brand that wants to show up when people ask Claude about its category should think about whether blocking ClaudeBot trades long-term training visibility for a short-term data-control preference, since AI answer visibility increasingly depends on what models have learned broadly. Using Crawl-delay instead of a full block is a middle option for sites mainly worried about server load rather than data use.

## How to block it

Add this to robots.txt to block ClaudeBot entirely:

User-agent: ClaudeBot
Disallow: /

Anthropic also supports a Crawl-delay directive if the goal is to slow the crawl rather than block it outright.

## FAQ

### Is ClaudeBot the same as the bot that lets Claude open a page a user links to?

No. That's Claude-User, a separate bot. ClaudeBot is specifically for collecting training data and is controlled independently.

### Does ClaudeBot respect robots.txt?

Yes. Anthropic says its bots honor standard "do not crawl" directives in robots.txt, and also support the non-standard Crawl-delay extension.

### Can I slow ClaudeBot down instead of blocking it completely?

Yes, using the Crawl-delay directive in robots.txt, which limits crawl rate without fully disallowing the bot.

## Sources

- ClaudeBot collects web content that could contribute to training Anthropic's generative AI models. Source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Anthropic's bots honor standard robots.txt directives. Source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- Anthropic supports the non-standard Crawl-delay directive to limit crawl rate. Source: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
