# Diffbot

## Quick answer

Diffbot is a web crawler run by Diffbot, a company that builds a structured knowledge graph and web search tools from crawled pages. It is a data and search infrastructure crawler, not a bot tied to a major consumer AI assistant.

## What it does

Diffbot crawls publicly available web pages proactively to build its Knowledge Graph and support its web search products, extracting structured facts from pages rather than just indexing raw text. By default, Diffbot's crawls follow a site's robots.txt instructions, including disallow rules and crawl-delay directives, the same as most compliant crawlers. Diffbot states that this general, proactive crawling is meant for building a general search engine and is not used for AI training, drawing a clear line between its own purpose and the training crawlers run by AI labs. In some partnership arrangements, Diffbot notes it may still crawl a site under a separate agreement with that site, which would operate outside the default robots.txt based crawling described here.

## Facts

- Operator: Diffbot
- Robots token: Diffbot
- User agent string(s): Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com/our-apis/crawler/)
- Purpose: search-index
- Respects robots.txt: yes

## Should you block it?

Blocking Diffbot mainly affects whether a site's pages contribute to Diffbot's Knowledge Graph and its own search product, rather than visibility in a well known consumer AI assistant, since Diffbot describes this crawling as separate from AI training. A business that relies on structured web data products might want to stay crawlable so its public information is represented accurately elsewhere. A publisher mainly concerned about a third party extracting and reselling structured data from its pages may prefer to disallow Diffbot, since by default it respects that choice through robots.txt. Because Diffbot says it does not use this crawl for AI training, the training-data concern that applies to crawlers like GPTBot or ClaudeBot does not apply here by default, though a specific partnership agreement could work differently.

## How to block it

Add this to robots.txt to block Diffbot:

User-agent: Diffbot
Disallow: /

Diffbot says its default crawls adhere to robots.txt instructions, including disallow and crawl-delay directives, so this rule should take effect for its standard crawling.

## FAQ

### Does Diffbot use crawled content to train AI models?

Diffbot describes its general, proactive crawling as being for building a search engine, and says it is not used for AI training.

### Does Diffbot respect robots.txt?

Yes, by default. Diffbot says its web crawls adhere to a site's robots.txt instructions, including disallow and crawl-delay directives.

### Can Diffbot crawl a site that disallows it?

In some partnership cases, Diffbot notes it may still crawl under a separate agreement with that specific site.

## Sources

- By default, Diffbot's web crawls adhere to a site's robots.txt instructions, including disallow and crawl-delay directives. Source: https://www.diffbot.com/docs/crawl/faq/robots-txt
- Diffbot's main crawler performs general, proactive web crawling for building a general search engine and is not used for AI training. Source: https://www.diffbot.com/docs/crawl/faq/robots-txt
