Skip to content
MarketHQ
AI crawler

MistralAI-Training

Mistral AI's ai model training crawler.

Quick answer

MistralAI-Training is Mistral AI's crawler for building the datasets behind training its own generative AI models. Mistral documents this bot specifically for that purpose, keeping it separate from any crawler it might use for search indexing or for fetching a page on a live user's behalf.

What it does

MistralAI-Training crawls publicly available web content specifically to help build the datasets Mistral uses when training its generative AI models. Mistral's own documentation describes this narrowly, naming training as the bot's purpose rather than bundling it with search indexing or other functions the way some other companies combine multiple jobs into one crawler. It follows robots.txt directives, so a site that disallows MistralAI-Training is opting that content out of Mistral's future training data collection going forward, though content already used in a completed training run isn't addressed by this documentation. Mistral doesn't publish a dedicated IP range file for this bot in this data, so verifying genuine traffic relies on checking the published user agent string instead.

Facts

Operator

Mistral AI

Robots token

MistralAI-Training

User agent string(s)

  • Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)

Purpose

AI model training

Respects robots.txt

Yes

From Mistral AI's documentation, checked Oct 2026.

Should you block it?

Disallowing MistralAI-Training keeps a site's future content out of the datasets Mistral uses to train its generative AI models, which matters to a publisher protecting original work or content it specifically doesn't want reused for AI training. Because Mistral documents this bot's purpose narrowly as training, blocking it is a fairly clean opt-out decision compared to bots that combine training with search indexing or live-answer functions, since there's no secondary function being affected here. Mistral isn't among the AI answer engines that MarketHQ tracks for brand visibility, so this decision is closer to a data-rights and training-data question than an AI-answer-visibility one; a brand focused mainly on showing up in AI assistant answers has less riding on this specific bot than it does on crawlers tied to the assistants it actually tracks.

How to block it

Add this to robots.txt to block MistralAI-Training from crawling any page: User-agent: MistralAI-Training Disallow: / This only affects Mistral's training-data collection; Mistral documents a separate purpose for this bot apart from search indexing or user-triggered fetches.

Frequently asked questions

What does MistralAI-Training actually collect?

Mistral says this bot crawls web content to help build datasets for training Mistral's generative AI models, a purpose it documents specifically for this bot.

Does blocking MistralAI-Training remove content already used in training?

Mistral's published documentation doesn't address that specifically, so it's safest to assume a robots.txt change affects future crawling, not a past completed training run.

Is MistralAI-Training the same bot used for search or live fetches?

No. Mistral documents MistralAI-Training specifically for training purposes, separate from any crawler it might use for other functions.

Related

  • AI crawler directory — every AI crawler's robots.txt token, purpose and facts in one place.
  • Free AI crawler checker — check which AI crawlers a domain's robots.txt actually allows or blocks.
  • Glossary — definitions of the AI-visibility terms that come up alongside crawler behavior.

Sources

  • MistralAI-Training crawls web content to help build datasets for training Mistral generative AI models. Source: https://docs.mistral.ai/en/robots (checked Oct 2026)

MarketHQ tracks brand mentions across communities, news, blogs, social and AI answers, and turns them into gap analysis and action plans.