Mistral is the open-weight model family built by Mistral AI, a
Paris lab that's spent three years positioning itself as Europe's
answer to OpenAI: permissively licensed, self-hostable, and pitched
on data sovereignty as much as raw benchmark scores. Its licensing
has actually swung back and forth over that time, tightening and
then loosening again, which matters more here than with most
families if you're planning a deployment around it. Here's who
builds it, how it's trained, and when it's the right call.
Who builds Mistral
Mistral AI was founded in April 2023 in Paris by
three co-founders who met at École Polytechnique: Arthur
Mensch (CEO), a former Google DeepMind researcher who
worked on Chinchilla; Guillaume Lample (Chief
Scientist), ex-Meta AI and one of the original creators of LLaMA;
and Timothée Lacroix (CTO), also ex-Meta AI. The
company raised a €105M seed round within weeks of founding and has
since raised well over €3B total, including a September 2025 round
led by chipmaker ASML that made ASML its largest shareholder at
roughly a €13B valuation. Mistral has leaned into sovereignty as a
selling point, landing a framework agreement with France's Ministry
of the Armed Forces and an expanded infrastructure partnership with
Microsoft Azure, while also shipping some of the most permissively
licensed frontier-adjacent open weights available.
A steady cadence of specialist releases, then a return to a fully open-weight flagship in late 2025.
How it's trained
Mistral's earliest models set the pattern the family still
follows: Mistral 7B (2023) was a dense 7.3B model
using grouped-query attention and sliding-window attention to
extend its effective context cheaply, and Mixtral
reused that same architecture as a sparse mixture-of-experts,
routing each token to 2 of 8 experts per layer, with an 8x22B
version reaching 141B total parameters at roughly 39B active.
Mistral Large 2 (2024) went back to dense, at 123B
parameters and 128K context, and was the family's strongest model
for a stretch, though it also marked the point where Mistral pulled
back from Apache 2.0 to a non-commercial research license for its
flagship weights.
Mistral's specialist lines each target a different capability
through targeted training rather than a bigger base model.
Magistral, the reasoning line launched in 2025,
was trained with a reinforcement learning pipeline built from
scratch around Group Relative Policy Optimization, notably without
distilling from another reasoning model's outputs and without a
separate critic model; Mistral's own technical report credits RL
alone with roughly a 50% jump in AIME math benchmark performance
for Magistral Medium. Devstral targets agentic
coding and SWE-bench-style tasks specifically, distinct from
Codestral's classic code-completion focus.
Pixtral adds a dedicated vision encoder on top of
the Mistral Large 2 decoder for native multimodal understanding,
and Voxtral does the same for speech.
Mistral Large 3, released December 2025, is
the family's biggest architectural jump: a granular
mixture-of-experts multimodal model at 675B total parameters with
only about 41B active per token, trained from scratch on 3,000
Nvidia H200 GPUs, with a 256K context window and native structured
function calling built in rather than bolted on. It's also where
Mistral's licensing swung back, Large 3 shipped under Apache 2.0,
reversing the more restrictive terms Large 2 had used the year
before.
COST-TO-PERFORMANCE
Mistral consistently markets active-parameter efficiency, aiming for near-frontier quality at a fraction of the active-parameter cost of comparably capable dense models.
FUNCTION CALLING
Native, structured function calling and JSON output are treated as first-class capabilities across the current generation, making the models a common choice for building agents.
DATA SOVEREIGNTY
Self-hostable open weights are core to Mistral's pitch to European and regulated customers who need inference to stay inside their own infrastructure, not a vendor's cloud.
When to use it
Mistral's lineup maps to a fairly clean set of tiers. Ministral
3, in 3B/8B/14B sizes, targets edge and cost-sensitive deployments.
Mistral Small, now unified in Small 4 (March 2026)
to cover reasoning, vision, and agentic coding in one checkpoint,
is the balanced general-purpose pick: a 119B-total mixture-of-experts
model with 128 experts and only 6B active per token, giving it a
40% latency reduction and 3x the throughput of Small 3 while
matching GPT-OSS 120B on reasoning benchmarks with shorter outputs.
It's still light enough to serve on a single high-end GPU or a
small multi-GPU node. Mistral Medium (3.5,
128B) targets longer-horizon cloud agentic and coding workloads
without the cost of the full flagship. Mistral Large 3 is the
option when you need genuinely frontier-adjacent quality and can
serve a 675B-parameter MoE, which in practice means a real
multi-GPU cluster, not a single card, even though only 41B
parameters activate per token.
Reach for Mistral specifically when native function calling, a
compliance or sovereignty requirement around self-hosting, or one
of its specialist variants, Magistral for reasoning, Devstral for
agentic coding, Pixtral or Voxtral for vision and speech, matches
your task. On raw benchmark leaderboards, 2026 commentary generally
places Mistral a tier below the very top open-weight models like
Kimi K3, DeepSeek, and Qwen3.5 on pure reasoning and agentic
coding, so if the entire decision comes down to squeezing out the
last few points on a benchmark, it's worth comparing against those
directly before committing.
Size tier tracks deployment fit, from a single edge device up to a served multi-GPU MoE cluster.
How to use it
Licensing is the one thing to check per-release rather than
assume with Mistral, more so than with most families. Mistral 7B
and Mixtral launched Apache 2.0; Mistral Large 2, Pixtral Large,
and some 2024 Ministral checkpoints shipped under the more
restrictive Mistral Research License, non-commercial without a
separate paid commercial license. By late 2025 the flagship line
reversed course: Mistral Large 3, Ministral 3, and Mistral
Small 4 all ship under Apache 2.0 again.
Don't assume a given checkpoint's license from an earlier release
in the same family, read the model card.
Weights are published on Hugging Face under the
mistralai org, in BF16, FP8, and NVFP4 variants, and
are also available through Amazon Bedrock, Azure AI Foundry, and
IBM watsonx for teams that want a managed path. For serving,
vLLM is Mistral's officially documented path,
including specific configs for Large 3's tensor-parallel
deployment across 8 H200 or B200 GPUs and a required
--tokenizer_mode mistral flag; Small 4 is far lighter
to serve, Mistral's own guidance puts it at 4x H100, 2x H200, or a
single B200 at minimum, and it's also distributed as an NVIDIA NIM
for production deployment. Mistral also maintains its own
lightweight mistral-inference reference implementation
and an official mistral-finetune repo for fine-tuning.
For quick evaluation before self-hosting,
Mistral's own La Plateforme offers a hosted API
with a documented migration path to self-hosted vLLM once you're
ready to move off it.
Where this fits at Numerata
Whichever Mistral tier fits your task, P95
is where you fine-tune it on your own data, on infrastructure you
control, private cloud or fully air-gapped, which pairs naturally
with Mistral's own sovereignty pitch rather than working against
it. NinetyFive serves it afterward,
turning whichever size tier you land on into production latency
instead of the naive multi-GPU serving times a model like Large 3
can otherwise produce.