Skip to content
MarketHQ
AI crawler

Googlebot

Google's search / answer indexing crawler.

Quick answer

Googlebot is Google's main crawler, the one responsible for discovering and fetching pages that go into Google Search's index. It's the crawler most site owners think of first when they consider blocking a bot, and it follows robots.txt rules automatically across the entire site.

What it does

Googlebot crawls the web to find and collect information that goes into building Google's search indexes, and Google also uses it for other product-specific crawls and analysis beyond core Search. It's grouped among Google's common crawlers, a category Google says always obeys robots.txt rules when crawling automatically, which is different from some of Google's special-case crawlers that can ignore a blanket rule under specific conditions. Googlebot identifies itself with a user agent string that includes 'Googlebot' along with a browser-like Chrome signature, and Google publishes a dedicated IP range file for it so site owners can verify genuine Googlebot traffic. Because Googlebot's crawl feeds Google's search index directly, how a site treats this one bot has an outsized effect on its visibility in Google Search specifically.

Facts

Operator

Google

Robots token

Googlebot

User agent string(s)

  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36

Purpose

Search / answer indexing

Respects robots.txt

Yes

IP ranges

https://developers.google.com/static/crawling/ipranges/common-crawlers.json

From Google's documentation, checked Oct 2026.

Should you block it?

Disallowing Googlebot is a significant decision for most sites, since Google's search indexes are built directly from what this crawler finds; blocking it stops new crawling right away, and pages already indexed will gradually drop out of Google Search results as Google notices they're no longer accessible. For a brand focused on being found through organic search, keeping Googlebot allowed is close to a baseline requirement rather than an optional choice, since there's no comparable substitute crawler that feeds the same index. The more nuanced decisions tend to be about which specific sections to disallow, like admin areas or duplicate content, rather than whether to block Googlebot site-wide. Because Googlebot's documented purpose is search indexing rather than AI model training, this decision is separate from the AI-training questions that apply to bots like GPTBot or Applebot-Extended.

How to block it

Add this to robots.txt to block Googlebot from crawling any page: user-agent: Googlebot disallow: / Most sites only disallow specific paths, like admin or duplicate-content sections, rather than blocking Googlebot site-wide, since that would remove the site from Google Search over time.

Frequently asked questions

What does Googlebot actually do?

Google says its common crawlers, including Googlebot, find information for building Google's search indexes and also perform other product-specific crawls and analysis.

Does Googlebot respect robots.txt?

Yes. Google says its common crawlers always obey robots.txt rules when crawling automatically, which includes Googlebot.

What happens if I block Googlebot site-wide?

Google stops crawling the site going forward, and previously indexed pages gradually drop out of Google Search results as Google notices they're no longer reachable.

Related

  • AI crawler directory — every AI crawler's robots.txt token, purpose and facts in one place.
  • Free AI crawler checker — check which AI crawlers a domain's robots.txt actually allows or blocks.
  • Glossary — definitions of the AI-visibility terms that come up alongside crawler behavior.

Sources

  • Google's common crawlers, including Googlebot, always obey robots.txt rules when crawling automatically. Source: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers (checked Oct 2026)
  • Google's common crawlers are used to build Google's search indexes and for other product-specific crawls and analysis. Source: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers (checked Oct 2026)

MarketHQ tracks brand mentions across communities, news, blogs, social and AI answers, and turns them into gap analysis and action plans.