BLOG · SLM SERIES

Mistral Explained

Mistral is the open-weight model family built by Mistral AI, a Paris lab that's spent three years positioning itself as Europe's answer to OpenAI: permissively licensed, self-hostable, and pitched on data sovereignty as much as raw benchmark scores. Its licensing has actually swung back and forth over that time, tightening and then loosening again, which matters more here than with most families if you're planning a deployment around it. Here's who builds it, how it's trained, and when it's the right call.

Who builds Mistral

Mistral AI was founded in April 2023 in Paris by three co-founders who met at École Polytechnique: Arthur Mensch (CEO), a former Google DeepMind researcher who worked on Chinchilla; Guillaume Lample (Chief Scientist), ex-Meta AI and one of the original creators of LLaMA; and Timothée Lacroix (CTO), also ex-Meta AI. The company raised a €105M seed round within weeks of founding and has since raised well over €3B total, including a September 2025 round led by chipmaker ASML that made ASML its largest shareholder at roughly a €13B valuation. Mistral has leaned into sovereignty as a selling point, landing a framework agreement with France's Ministry of the Armed Forces and an expanded infrastructure partnership with Microsoft Azure, while also shipping some of the most permissively licensed frontier-adjacent open weights available.

Mistral release timeline A timeline from Mistral 7B in September 2023 through Mixtral, Mistral Large 2, the Magistral and Devstral specialist lines, to Mistral Large 3 in December 2025 as the current flagship. 7B / MIXTRAL 2023 LARGE 2 2024 MAGISTRAL / DEVSTRAL 2025 LARGE 3 DEC 2025
A steady cadence of specialist releases, then a return to a fully open-weight flagship in late 2025.

How it's trained

Mistral's earliest models set the pattern the family still follows: Mistral 7B (2023) was a dense 7.3B model using grouped-query attention and sliding-window attention to extend its effective context cheaply, and Mixtral reused that same architecture as a sparse mixture-of-experts, routing each token to 2 of 8 experts per layer, with an 8x22B version reaching 141B total parameters at roughly 39B active. Mistral Large 2 (2024) went back to dense, at 123B parameters and 128K context, and was the family's strongest model for a stretch, though it also marked the point where Mistral pulled back from Apache 2.0 to a non-commercial research license for its flagship weights.

Mistral's specialist lines each target a different capability through targeted training rather than a bigger base model. Magistral, the reasoning line launched in 2025, was trained with a reinforcement learning pipeline built from scratch around Group Relative Policy Optimization, notably without distilling from another reasoning model's outputs and without a separate critic model; Mistral's own technical report credits RL alone with roughly a 50% jump in AIME math benchmark performance for Magistral Medium. Devstral targets agentic coding and SWE-bench-style tasks specifically, distinct from Codestral's classic code-completion focus. Pixtral adds a dedicated vision encoder on top of the Mistral Large 2 decoder for native multimodal understanding, and Voxtral does the same for speech.

Mistral Large 3, released December 2025, is the family's biggest architectural jump: a granular mixture-of-experts multimodal model at 675B total parameters with only about 41B active per token, trained from scratch on 3,000 Nvidia H200 GPUs, with a 256K context window and native structured function calling built in rather than bolted on. It's also where Mistral's licensing swung back, Large 3 shipped under Apache 2.0, reversing the more restrictive terms Large 2 had used the year before.

COST-TO-PERFORMANCE

Mistral consistently markets active-parameter efficiency, aiming for near-frontier quality at a fraction of the active-parameter cost of comparably capable dense models.

FUNCTION CALLING

Native, structured function calling and JSON output are treated as first-class capabilities across the current generation, making the models a common choice for building agents.

DATA SOVEREIGNTY

Self-hostable open weights are core to Mistral's pitch to European and regulated customers who need inference to stay inside their own infrastructure, not a vendor's cloud.

When to use it

Mistral's lineup maps to a fairly clean set of tiers. Ministral 3, in 3B/8B/14B sizes, targets edge and cost-sensitive deployments. Mistral Small, now unified in Small 4 (March 2026) to cover reasoning, vision, and agentic coding in one checkpoint, is the balanced general-purpose pick: a 119B-total mixture-of-experts model with 128 experts and only 6B active per token, giving it a 40% latency reduction and 3x the throughput of Small 3 while matching GPT-OSS 120B on reasoning benchmarks with shorter outputs. It's still light enough to serve on a single high-end GPU or a small multi-GPU node. Mistral Medium (3.5, 128B) targets longer-horizon cloud agentic and coding workloads without the cost of the full flagship. Mistral Large 3 is the option when you need genuinely frontier-adjacent quality and can serve a 675B-parameter MoE, which in practice means a real multi-GPU cluster, not a single card, even though only 41B parameters activate per token.

Reach for Mistral specifically when native function calling, a compliance or sovereignty requirement around self-hosting, or one of its specialist variants, Magistral for reasoning, Devstral for agentic coding, Pixtral or Voxtral for vision and speech, matches your task. On raw benchmark leaderboards, 2026 commentary generally places Mistral a tier below the very top open-weight models like Kimi K3, DeepSeek, and Qwen3.5 on pure reasoning and agentic coding, so if the entire decision comes down to squeezing out the last few points on a benchmark, it's worth comparing against those directly before committing.

Mistral size tiers mapped to deployment fit A chart mapping Mistral model sizes to deployment scenarios: Ministral 3 in 3B to 14B for edge use, Mistral Small 4 at 119B total with 6B active parameters for balanced single-GPU serving with reasoning, vision, and coding unified, Mistral Medium 3.5 at 128B for cloud agentic workloads, and Mistral Large 3 at 675B total with 41B active parameters as the flagship mixture-of-experts model. 3B-14B EDGE (MINISTRAL 3) 119B-A6B SMALL 4, ALL-IN-ONE 128B MEDIUM 3.5, CLOUD AGENTS 675B-A41B (LARGE 3) FLAGSHIP
Size tier tracks deployment fit, from a single edge device up to a served multi-GPU MoE cluster.

How to use it

Licensing is the one thing to check per-release rather than assume with Mistral, more so than with most families. Mistral 7B and Mixtral launched Apache 2.0; Mistral Large 2, Pixtral Large, and some 2024 Ministral checkpoints shipped under the more restrictive Mistral Research License, non-commercial without a separate paid commercial license. By late 2025 the flagship line reversed course: Mistral Large 3, Ministral 3, and Mistral Small 4 all ship under Apache 2.0 again. Don't assume a given checkpoint's license from an earlier release in the same family, read the model card.

Weights are published on Hugging Face under the mistralai org, in BF16, FP8, and NVFP4 variants, and are also available through Amazon Bedrock, Azure AI Foundry, and IBM watsonx for teams that want a managed path. For serving, vLLM is Mistral's officially documented path, including specific configs for Large 3's tensor-parallel deployment across 8 H200 or B200 GPUs and a required --tokenizer_mode mistral flag; Small 4 is far lighter to serve, Mistral's own guidance puts it at 4x H100, 2x H200, or a single B200 at minimum, and it's also distributed as an NVIDIA NIM for production deployment. Mistral also maintains its own lightweight mistral-inference reference implementation and an official mistral-finetune repo for fine-tuning. For quick evaluation before self-hosting, Mistral's own La Plateforme offers a hosted API with a documented migration path to self-hosted vLLM once you're ready to move off it.

Where this fits at Numerata

Whichever Mistral tier fits your task, P95 is where you fine-tune it on your own data, on infrastructure you control, private cloud or fully air-gapped, which pairs naturally with Mistral's own sovereignty pitch rather than working against it. NinetyFive serves it afterward, turning whichever size tier you land on into production latency instead of the naive multi-GPU serving times a model like Large 3 can otherwise produce.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog