Skip to content
MarketHQ
AI crawler

Meta-ExternalAgent

Meta's ai model training crawler.

Quick answer

Meta-ExternalAgent is Meta's crawler for collecting web content that can be used for training foundation AI models or for improving Meta's products by indexing content directly. It's a different bot from Meta-WebIndexer, which is focused specifically on the search-style indexing that feeds Meta AI's answers.

What it does

Meta-ExternalAgent crawls publicly available web pages for use cases Meta describes specifically as training foundation AI models or improving its products through direct content indexing. That combined purpose, training and product indexing, makes it broader than a bot built for just one of those jobs; Meta documents it separately from Meta-WebIndexer, which focuses on indexing for Meta AI's search-style answers. Meta-ExternalAgent follows robots.txt rules, and a site can write a rule that allows most content while disallowing specific sections, like a private directory, if it wants the crawl to continue for public content while excluding gated or sensitive paths. Meta doesn't publish a separate IP range file specifically for Meta-ExternalAgent in this data.

Facts

Operator

Meta

Robots token

meta-externalagent

User agent string(s)

  • meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers)
  • meta-externalagent/1.1

Purpose

AI model training

Respects robots.txt

Yes

From Meta's documentation, checked Oct 2026.

Should you block it?

Disallowing Meta-ExternalAgent opts a site out of having its content used both for training Meta's foundation AI models and for the direct product-indexing use Meta also documents for this bot, so it's a broader opt-out than blocking a crawler used for only one of those purposes. A publisher protecting original reporting or content it doesn't want reused for AI training would have reason to block it, while accepting that this also removes the direct-indexing benefit Meta describes. Because Meta-ExternalAgent's purpose is training and indexing, not live user-triggered fetches, blocking it doesn't affect a real-time Meta AI answer to a specific user's question the way blocking a user-fetch bot would; the tradeoff here is mainly about long-term training data and whether a site wants its content folded into Meta's AI products at all.

How to block it

Add this to robots.txt to block Meta-ExternalAgent from crawling any page, while still allowing public sections if desired: User-agent: meta-externalagent Disallow: / Use 'Allow: /' with specific 'Disallow:' lines for private sections to permit crawling of public content only.

Frequently asked questions

What does Meta-ExternalAgent use crawled content for?

Meta documents it as crawling for use cases such as training foundation AI models or improving products by indexing content directly.

Is Meta-ExternalAgent the same as Meta-WebIndexer?

No. Meta-WebIndexer is documented separately, focused on search-style indexing for Meta AI. Meta-ExternalAgent covers training and direct product indexing.

Does Meta-ExternalAgent follow robots.txt?

Yes, based on this bot's documented robots.txt behavior. A site can write rules allowing public sections while disallowing specific private paths.

Related

  • AI crawler directory — every AI crawler's robots.txt token, purpose and facts in one place.
  • Free AI crawler checker — check which AI crawlers a domain's robots.txt actually allows or blocks.
  • Glossary — definitions of the AI-visibility terms that come up alongside crawler behavior.

Sources

  • Meta-ExternalAgent crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. Source: https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers (checked Oct 2026)

MarketHQ tracks brand mentions across communities, news, blogs, social and AI answers, and turns them into gap analysis and action plans.