Diffbot
Diffbot's search / answer indexing crawler.
Quick answer
Diffbot is a web crawler run by Diffbot, a company that builds a structured knowledge graph and web search tools from crawled pages. It is a data and search infrastructure crawler, not a bot tied to a major consumer AI assistant.
What it does
Diffbot crawls publicly available web pages proactively to build its Knowledge Graph and support its web search products, extracting structured facts from pages rather than just indexing raw text. By default, Diffbot's crawls follow a site's robots.txt instructions, including disallow rules and crawl-delay directives, the same as most compliant crawlers. Diffbot states that this general, proactive crawling is meant for building a general search engine and is not used for AI training, drawing a clear line between its own purpose and the training crawlers run by AI labs. In some partnership arrangements, Diffbot notes it may still crawl a site under a separate agreement with that site, which would operate outside the default robots.txt based crawling described here.
Facts
Operator
Diffbot
Robots token
Diffbot
User agent string(s)
- Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com/our-apis/crawler/)
Purpose
Search / answer indexing
Respects robots.txt
Yes
From Diffbot's documentation, checked Oct 2026.
Should you block it?
Blocking Diffbot mainly affects whether a site's pages contribute to Diffbot's Knowledge Graph and its own search product, rather than visibility in a well known consumer AI assistant, since Diffbot describes this crawling as separate from AI training. A business that relies on structured web data products might want to stay crawlable so its public information is represented accurately elsewhere. A publisher mainly concerned about a third party extracting and reselling structured data from its pages may prefer to disallow Diffbot, since by default it respects that choice through robots.txt. Because Diffbot says it does not use this crawl for AI training, the training-data concern that applies to crawlers like GPTBot or ClaudeBot does not apply here by default, though a specific partnership agreement could work differently.
How to block it
Add this to robots.txt to block Diffbot: User-agent: Diffbot Disallow: / Diffbot says its default crawls adhere to robots.txt instructions, including disallow and crawl-delay directives, so this rule should take effect for its standard crawling.
Frequently asked questions
Does Diffbot use crawled content to train AI models?
Diffbot describes its general, proactive crawling as being for building a search engine, and says it is not used for AI training.
Does Diffbot respect robots.txt?
Yes, by default. Diffbot says its web crawls adhere to a site's robots.txt instructions, including disallow and crawl-delay directives.
Can Diffbot crawl a site that disallows it?
In some partnership cases, Diffbot notes it may still crawl under a separate agreement with that specific site.
Related
- AI crawler directory — every AI crawler's robots.txt token, purpose and facts in one place.
- Free AI crawler checker — check which AI crawlers a domain's robots.txt actually allows or blocks.
- Glossary — definitions of the AI-visibility terms that come up alongside crawler behavior.
Sources
- By default, Diffbot's web crawls adhere to a site's robots.txt instructions, including disallow and crawl-delay directives. Source: https://www.diffbot.com/docs/crawl/faq/robots-txt (checked Oct 2026)
- Diffbot's main crawler performs general, proactive web crawling for building a general search engine and is not used for AI training. Source: https://www.diffbot.com/docs/crawl/faq/robots-txt (checked Oct 2026)
MarketHQ tracks brand mentions across communities, news, blogs, social and AI answers, and turns them into gap analysis and action plans.