Google-Extended
Google's ai model training crawler.
Quick answer
Google-Extended is a robots.txt control, not a separate crawler with its own traffic. It lets a site decide whether content Google already crawls can also be used to train future Gemini models or ground Gemini's answers with current web information.
What it does
Google-Extended doesn't send its own requests; instead, it is a standalone token publishers add to robots.txt to control a specific use of content Google's existing crawlers already gather. Setting a Google-Extended rule tells Google whether that crawled content may be used to train future Gemini models or to ground Gemini's answers with up to date information, separate from how that same content is used for Google Search. Google states clearly that Google-Extended has no impact on a site's inclusion in Google Search and is not used as a ranking signal, so disallowing it does not touch regular search visibility. Because it piggybacks on Google's common crawlers, Google-Extended follows the published IP ranges for those crawlers rather than having a separate range of its own. MarketHQ tracks how brands appear in Gemini's answers as part of its wider AI visibility tracking across ChatGPT, Claude, Gemini, Perplexity, Grok, and DeepSeek.
Facts
Operator
Robots token
Google-Extended
User agent string(s)
- (no separate HTTP user-agent string; crawls with existing Google user agents, controlled via the Google-Extended robots.txt token)
Purpose
AI model training
Respects robots.txt
Yes
IP ranges
https://developers.google.com/static/crawling/ipranges/common-crawlers.json
From Google's documentation, checked Oct 2026.
Should you block it?
Disallowing Google-Extended keeps a site's content out of future Gemini training and out of the material Gemini uses to ground its answers, which appeals to a publisher that wants tighter control over how its work feeds AI models. The tradeoff is that Gemini is one of the AI assistants people increasingly ask for recommendations and comparisons, so opting out of this training and grounding data can make a brand less likely to be accurately represented when someone asks Gemini about its category. Google is explicit that none of this affects a site's regular Google Search ranking or inclusion, so a site does not have to weigh search visibility against this choice; it is purely about Gemini. A brand actively trying to show up well in AI answers, including through tools like MarketHQ that track AI visibility, generally benefits from leaving Google-Extended allowed so Gemini has current information to draw from.
How to block it
Add this to robots.txt to opt out of Google-Extended: user-agent: Google-Extended disallow: / This only affects training and grounding data for future Gemini models; it has no effect on Google Search inclusion or ranking.
Frequently asked questions
Does Google-Extended affect my Google Search ranking?
No. Google says Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal.
What does Google-Extended actually control?
It lets a site control whether Google may use crawled content to train future Gemini models or ground Gemini's answers.
Does Google-Extended send its own crawl requests?
No. It is a robots.txt token that controls use of content gathered by Google's existing crawlers, not a separate bot.
Related
- AI crawler directory — every AI crawler's robots.txt token, purpose and facts in one place.
- Free AI crawler checker — check which AI crawlers a domain's robots.txt actually allows or blocks.
- Glossary — definitions of the AI-visibility terms that come up alongside crawler behavior.
Sources
- Google-Extended lets publishers control whether Google may use crawled content to train future Gemini models and for grounding. Source: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers (checked Oct 2026)
- Google-Extended does not impact a site's inclusion in Google Search or act as a ranking signal. Source: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers (checked Oct 2026)
MarketHQ tracks brand mentions across communities, news, blogs, social and AI answers, and turns them into gap analysis and action plans.