Llama is Meta's open-weight model family, and for most of 2023
and 2024 it was the default choice for anyone building on open
weights, less because it topped every benchmark and more because
the ecosystem around it, fine-tuning tools, serving frameworks,
community checkpoints, was unmatched. That's still mostly true. What's
changed is the model quality story: Llama 4 landed to a mixed
reception in 2025, and by 2026 several Chinese open-weight labs
have taken the benchmark lead. Here's who builds it, how it's
trained, and an honest read on where it fits today.
Who builds Llama
Llama started inside Meta AI's FAIR (Fundamental
AI Research) group. In June 2025, Meta restructured around it,
forming Meta Superintelligence Labs, folding in
FAIR and Meta's other AI teams, after investing $14.3 billion in
Scale AI and hiring its CEO, Alexandr Wang, as
Meta's first Chief AI Officer to lead the new group. The
reorganization brought real turbulence: roughly 600 layoffs hit
the new lab that October, and in November 2025 Yann
LeCun, FAIR's co-founder and Meta's chief AI scientist for
over a decade, left to start his own lab, reportedly over
disagreements about prioritizing large language models over his
own research direction. Release-wise, Llama has moved from Llama 1
(Feb 2023, research-only) through Llama 2 (2023), Llama 3 and 3.1's
405B flagship (2024), the multimodal Llama 3.2 (2024), Llama 3.3
(Dec 2024), to Llama 4's Scout and Maverick (April 2025). A larger
Llama 4 Behemoth model was previewed alongside them but never
publicly released, and by mid-2026 reporting suggested its training
had stalled. In April 2026, Meta shipped a separate closed-weight
model, Muse Spark, its first major model not released as open
weights, a signal worth watching if you're planning around Llama's
continued openness rather than assuming it.
A steady yearly cadence through 2024, then a shift to mixture-of-experts with Llama 4 in 2025.
How it's trained
Through Llama 3, the family was straightforwardly
dense. Llama 1 (2023) trained on roughly 1-1.4
trillion tokens at up to 65B parameters with a 2K context window;
Llama 2 doubled the pretraining data to 2 trillion tokens and
extended context to 4K; Llama 3 (April 2024) jumped pretraining to
roughly 15 trillion tokens with an 8K window, and its July 2024
follow-up, Llama 3.1, extended context to 128K and added a 405B
dense flagship, the largest openly released dense model of its
generation. Llama 3 also moved the tokenizer from Llama 2's 32K
SentencePiece vocabulary to a 128K-token, tiktoken-based BPE
vocabulary for better multilingual efficiency. Llama 3.2 (September
2024) added vision to the 11B and 90B sizes through a separate
image adapter bolted onto the existing text backbone, rather than
training a natively multimodal model from scratch, alongside 1B
and 3B text-only models aimed squarely at on-device use.
Llama 4 (April 2025) is the family's first
mixture-of-experts generation and its first
attempt at native multimodality, fusing text, image, and video
understanding from early in training rather than adapting a
finished text model afterward. Scout is 109B total parameters with
17B active, and pushed context all the way to 10 million tokens;
Maverick is larger at 400B total with the same 17B active, at a
1-million-token context window. A third, larger Behemoth model,
reportedly around 2 trillion total parameters, was meant to serve
as a teacher model that Scout and Maverick distilled from, but
Meta never released it publicly, and by mid-2026 its training was
reported to have stalled on MoE-routing and attention-scaling
issues at that size. Post-training across the family combines
supervised fine-tuning with RLHF and rejection sampling, with
Llama 3 onward adding DPO stages on top.
ECOSYSTEM MATURITY
The largest fine-tuning and serving ecosystem of any open-weight family: more community checkpoints, more first-tested framework support, and the deepest enterprise cloud integration.
ON-DEVICE TIER
The 1B and 3B Llama 3.2 models were built specifically for on-device deployment, with early partnerships from Qualcomm, MediaTek, and Arm for mobile and edge hardware.
LONG CONTEXT
Llama 4 Scout's 10-million-token context window is among the longest publicly available at any size, useful for tasks that need to reason over very large documents or codebases at once.
When to use it
Size tiers still map cleanly: 1B/3B for edge and on-device,
8B/70B as the general-purpose server-side default that most
existing Llama tooling targets, 405B (Llama 3.1) as a dense
frontier option, and Scout or Maverick when you specifically need
Llama 4's long context or native multimodality. Where the picture
has shifted is quality relative to the field. Llama 4 was widely
described as a disappointing generation on release, and
independent 2026 leaderboards generally place Qwen, DeepSeek, and
Kimi ahead of Llama on raw coding, reasoning, and agentic
benchmarks, some by a wide margin. Llama hasn't fallen out of use,
it's still one of the most deployed open-weight families, but it's
no longer the default answer to "which open model is best," the
way it arguably was through 2024.
What still makes Llama worth reaching for isn't the raw
benchmark score, it's the ecosystem: more mature fine-tuning
tooling, wider first-party support across every serving framework
and cloud provider, and a much larger base of prior art if your
team is fine-tuning rather than building from scratch. If squeezing
out the best possible benchmark result is the priority and you're
comfortable with a thinner tooling ecosystem, it's worth comparing
Llama directly against Qwen,
Kimi, or DeepSeek
before committing.
Size tier tracks deployment fit, with Llama 4's MoE models reserved for long-context or multimodal tasks specifically.
How to use it
Llama's license is the thing most people get wrong about it:
despite Meta calling it open source, the Llama Community
License is not OSI-approved, and the Open Source
Initiative has publicly disputed Meta's framing. The practical
restriction that matters is the 700 million monthly
active user clause: if your product, or an affiliate's,
crosses that threshold, you have to request a separate license
from Meta, granted entirely at Meta's discretion, not automatically.
Below that bar, it's free to use commercially. The license also
requires attribution ("Built with Llama") and Llama-branding in
derivative products. One restriction did loosen over time: earlier
versions banned using Llama's outputs to improve any other large
language model outside the Llama family; starting with Llama 3.1,
that's allowed as long as you attribute Llama in the resulting
model and its documentation, which matters if distillation is part
of your plan.
Weights are distributed through Hugging Face under the gated
meta-llama org, requiring license acceptance before
download, and through Meta's own developer channels. Serving
support is Llama's clearest advantage: vLLM, Hugging Face
TGI, Ollama, and llama.cpp all treat it as a first-class
target, and it's natively available on every major managed
platform, AWS Bedrock, Azure AI Foundry, Google Vertex, Together,
Fireworks, Groq, and Databricks among them. The fine-tuning
ecosystem is similarly the deepest of any open-weight family:
more LoRA and QLoRA tooling has been built and tested against
Llama first than against any other model, and it has by far the
largest count of community fine-tunes published on Hugging Face,
even in a year where its base model quality isn't the headline.
Where this fits at Numerata
If Llama's ecosystem maturity is what makes it the right choice
for your team, P95 is where you
fine-tune it on your own data, on infrastructure you control,
private cloud or fully air-gapped, and NinetyFive
serves it afterward, so you keep Llama's tooling advantage without
giving up the latency or data control a managed API would cost
you.