# Meta-ExternalAgent

## Quick answer

Meta-ExternalAgent is Meta's crawler for collecting web content that can be used for training foundation AI models or for improving Meta's products by indexing content directly. It's a different bot from Meta-WebIndexer, which is focused specifically on the search-style indexing that feeds Meta AI's answers.

## What it does

Meta-ExternalAgent crawls publicly available web pages for use cases Meta describes specifically as training foundation AI models or improving its products through direct content indexing. That combined purpose, training and product indexing, makes it broader than a bot built for just one of those jobs; Meta documents it separately from Meta-WebIndexer, which focuses on indexing for Meta AI's search-style answers. Meta-ExternalAgent follows robots.txt rules, and a site can write a rule that allows most content while disallowing specific sections, like a private directory, if it wants the crawl to continue for public content while excluding gated or sensitive paths. Meta doesn't publish a separate IP range file specifically for Meta-ExternalAgent in this data.

## Facts

- Operator: Meta
- Robots token: meta-externalagent
- User agent string(s): meta-externalagent/1.1 (+/documentation/sharing/webmasters/web-crawlers); meta-externalagent/1.1
- Purpose: training
- Respects robots.txt: yes

## Should you block it?

Disallowing Meta-ExternalAgent opts a site out of having its content used both for training Meta's foundation AI models and for the direct product-indexing use Meta also documents for this bot, so it's a broader opt-out than blocking a crawler used for only one of those purposes. A publisher protecting original reporting or content it doesn't want reused for AI training would have reason to block it, while accepting that this also removes the direct-indexing benefit Meta describes. Because Meta-ExternalAgent's purpose is training and indexing, not live user-triggered fetches, blocking it doesn't affect a real-time Meta AI answer to a specific user's question the way blocking a user-fetch bot would; the tradeoff here is mainly about long-term training data and whether a site wants its content folded into Meta's AI products at all.

## How to block it

Add this to robots.txt to block Meta-ExternalAgent from crawling any page, while still allowing public sections if desired:

User-agent: meta-externalagent
Disallow: /

Use 'Allow: /' with specific 'Disallow:' lines for private sections to permit crawling of public content only.

## FAQ

### What does Meta-ExternalAgent use crawled content for?

Meta documents it as crawling for use cases such as training foundation AI models or improving products by indexing content directly.

### Is Meta-ExternalAgent the same as Meta-WebIndexer?

No. Meta-WebIndexer is documented separately, focused on search-style indexing for Meta AI. Meta-ExternalAgent covers training and direct product indexing.

### Does Meta-ExternalAgent follow robots.txt?

Yes, based on this bot's documented robots.txt behavior. A site can write rules allowing public sections while disallowing specific private paths.

## Sources

- Meta-ExternalAgent crawls the web for use cases such as training foundation AI models or improving products by indexing content directly. Source: https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers
