GPTBot
OpenAI's ai model training crawler.
Quick answer
GPTBot is OpenAI's web crawler, used to collect content that may train OpenAI's generative AI foundation models, the kind of models behind ChatGPT. It respects robots.txt and crawls sites that allow it, separate from the bots that power live ChatGPT search answers.
What it does
GPTBot crawls publicly available web pages to gather text that OpenAI may use when training its foundation models, with the stated goal of making those models more useful and safe. It's distinct from the bots OpenAI uses for real-time ChatGPT search results and ads, which run under separate names and separate robots.txt rules, so blocking GPTBot doesn't necessarily remove a site from live ChatGPT answers. GPTBot respects standard robots.txt directives, meaning a site that disallows it should not have new content crawled by this bot going forward, though previously crawled and already-trained-on content isn't retroactively removed. OpenAI publishes GPTBot's IP ranges as a JSON file so site owners can verify that traffic claiming to be GPTBot is genuine before deciding how to treat it.
Facts
Operator
OpenAI
Robots token
GPTBot
User agent string(s)
- Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
Purpose
AI model training
Respects robots.txt
Yes
IP ranges
https://openai.com/gptbot.json
From OpenAI's documentation, checked Oct 2026.
Should you block it?
Blocking GPTBot keeps a site's content out of future OpenAI model training, which some site owners want to protect original reporting, proprietary data, or content they monetize directly. The tradeoff is that training data and live-answer data aren't the same thing for every OpenAI product, so blocking GPTBot specifically targets training use rather than guaranteeing removal from ChatGPT's conversational answers, which may draw on other bots or cached data. A brand trying to appear in AI answers when people ask about its category should weigh that visibility goal against the training concern; blocking every AI crawler across the board can quietly reduce how often a brand gets mentioned when someone asks an AI assistant for a recommendation. Many sites choose a middle path: allow crawling of public marketing and product pages while disallowing bots on gated, proprietary, or paid content directories.
How to block it
Add this to the site's robots.txt file to block GPTBot from crawling any page: User-agent: GPTBot Disallow: / To allow it instead, use Allow: / or simply omit a rule for GPTBot, since allowing is the default when no rule is set.
Frequently asked questions
Does GPTBot power live ChatGPT search answers?
No. GPTBot is used for training OpenAI's models. Live ChatGPT search answers and ads are handled by separate OpenAI bots with their own robots.txt rules.
How can I verify traffic claiming to be GPTBot?
OpenAI publishes GPTBot's IP ranges as a JSON file at openai.com/gptbot.json, so incoming traffic can be checked against that list.
Does blocking GPTBot remove content already used in training?
No. Disallowing GPTBot stops future crawling for training purposes; it doesn't retroactively remove content a model may have already been trained on.
Related
- AI crawler directory — every AI crawler's robots.txt token, purpose and facts in one place.
- Free AI crawler checker — check which AI crawlers a domain's robots.txt actually allows or blocks.
- Glossary — definitions of the AI-visibility terms that come up alongside crawler behavior.
Sources
- GPTBot is used to crawl content that may be used in training OpenAI's generative AI foundation models. Source: https://developers.openai.com/api/docs/bots (checked Oct 2026)
- OpenAI publishes GPTBot's IP ranges as JSON so site owners can verify traffic. Source: https://developers.openai.com/api/docs/bots (checked Oct 2026)
MarketHQ tracks brand mentions across communities, news, blogs, social and AI answers, and turns them into gap analysis and action plans.