BLOG · SLM SERIES

Llama Explained

Llama is Meta's open-weight model family, and for most of 2023 and 2024 it was the default choice for anyone building on open weights, less because it topped every benchmark and more because the ecosystem around it, fine-tuning tools, serving frameworks, community checkpoints, was unmatched. That's still mostly true. What's changed is the model quality story: Llama 4 landed to a mixed reception in 2025, and by 2026 several Chinese open-weight labs have taken the benchmark lead. Here's who builds it, how it's trained, and an honest read on where it fits today.

Who builds Llama

Llama started inside Meta AI's FAIR (Fundamental AI Research) group. In June 2025, Meta restructured around it, forming Meta Superintelligence Labs, folding in FAIR and Meta's other AI teams, after investing $14.3 billion in Scale AI and hiring its CEO, Alexandr Wang, as Meta's first Chief AI Officer to lead the new group. The reorganization brought real turbulence: roughly 600 layoffs hit the new lab that October, and in November 2025 Yann LeCun, FAIR's co-founder and Meta's chief AI scientist for over a decade, left to start his own lab, reportedly over disagreements about prioritizing large language models over his own research direction. Release-wise, Llama has moved from Llama 1 (Feb 2023, research-only) through Llama 2 (2023), Llama 3 and 3.1's 405B flagship (2024), the multimodal Llama 3.2 (2024), Llama 3.3 (Dec 2024), to Llama 4's Scout and Maverick (April 2025). A larger Llama 4 Behemoth model was previewed alongside them but never publicly released, and by mid-2026 reporting suggested its training had stalled. In April 2026, Meta shipped a separate closed-weight model, Muse Spark, its first major model not released as open weights, a signal worth watching if you're planning around Llama's continued openness rather than assuming it.

Llama release timeline A timeline from Llama 1 in February 2023 through Llama 2, Llama 3 and its 405B flagship, the multimodal Llama 3.2, to Llama 4's Scout and Maverick in April 2025 as the current generation. LLAMA 1 2023 LLAMA 2 2023 LLAMA 3 / 3.1 / 3.2 2024 LLAMA 4 APRIL 2025
A steady yearly cadence through 2024, then a shift to mixture-of-experts with Llama 4 in 2025.

How it's trained

Through Llama 3, the family was straightforwardly dense. Llama 1 (2023) trained on roughly 1-1.4 trillion tokens at up to 65B parameters with a 2K context window; Llama 2 doubled the pretraining data to 2 trillion tokens and extended context to 4K; Llama 3 (April 2024) jumped pretraining to roughly 15 trillion tokens with an 8K window, and its July 2024 follow-up, Llama 3.1, extended context to 128K and added a 405B dense flagship, the largest openly released dense model of its generation. Llama 3 also moved the tokenizer from Llama 2's 32K SentencePiece vocabulary to a 128K-token, tiktoken-based BPE vocabulary for better multilingual efficiency. Llama 3.2 (September 2024) added vision to the 11B and 90B sizes through a separate image adapter bolted onto the existing text backbone, rather than training a natively multimodal model from scratch, alongside 1B and 3B text-only models aimed squarely at on-device use.

Llama 4 (April 2025) is the family's first mixture-of-experts generation and its first attempt at native multimodality, fusing text, image, and video understanding from early in training rather than adapting a finished text model afterward. Scout is 109B total parameters with 17B active, and pushed context all the way to 10 million tokens; Maverick is larger at 400B total with the same 17B active, at a 1-million-token context window. A third, larger Behemoth model, reportedly around 2 trillion total parameters, was meant to serve as a teacher model that Scout and Maverick distilled from, but Meta never released it publicly, and by mid-2026 its training was reported to have stalled on MoE-routing and attention-scaling issues at that size. Post-training across the family combines supervised fine-tuning with RLHF and rejection sampling, with Llama 3 onward adding DPO stages on top.

ECOSYSTEM MATURITY

The largest fine-tuning and serving ecosystem of any open-weight family: more community checkpoints, more first-tested framework support, and the deepest enterprise cloud integration.

ON-DEVICE TIER

The 1B and 3B Llama 3.2 models were built specifically for on-device deployment, with early partnerships from Qualcomm, MediaTek, and Arm for mobile and edge hardware.

LONG CONTEXT

Llama 4 Scout's 10-million-token context window is among the longest publicly available at any size, useful for tasks that need to reason over very large documents or codebases at once.

When to use it

Size tiers still map cleanly: 1B/3B for edge and on-device, 8B/70B as the general-purpose server-side default that most existing Llama tooling targets, 405B (Llama 3.1) as a dense frontier option, and Scout or Maverick when you specifically need Llama 4's long context or native multimodality. Where the picture has shifted is quality relative to the field. Llama 4 was widely described as a disappointing generation on release, and independent 2026 leaderboards generally place Qwen, DeepSeek, and Kimi ahead of Llama on raw coding, reasoning, and agentic benchmarks, some by a wide margin. Llama hasn't fallen out of use, it's still one of the most deployed open-weight families, but it's no longer the default answer to "which open model is best," the way it arguably was through 2024.

What still makes Llama worth reaching for isn't the raw benchmark score, it's the ecosystem: more mature fine-tuning tooling, wider first-party support across every serving framework and cloud provider, and a much larger base of prior art if your team is fine-tuning rather than building from scratch. If squeezing out the best possible benchmark result is the priority and you're comfortable with a thinner tooling ecosystem, it's worth comparing Llama directly against Qwen, Kimi, or DeepSeek before committing.

Llama size tiers mapped to deployment fit A chart mapping Llama model sizes to deployment scenarios: 1B and 3B for edge and on-device use, 8B and 70B as the general-purpose server-side default, 405B as a dense frontier option, and Llama 4 Scout or Maverick, mixture-of-experts models with 17 billion active parameters, for long context or native multimodal tasks. 1B-3B EDGE / ON-DEVICE 8B-70B GENERAL DEFAULT 405B DENSE FRONTIER SCOUT / MAVERICK LONG CONTEXT / MULTIMODAL
Size tier tracks deployment fit, with Llama 4's MoE models reserved for long-context or multimodal tasks specifically.

How to use it

Llama's license is the thing most people get wrong about it: despite Meta calling it open source, the Llama Community License is not OSI-approved, and the Open Source Initiative has publicly disputed Meta's framing. The practical restriction that matters is the 700 million monthly active user clause: if your product, or an affiliate's, crosses that threshold, you have to request a separate license from Meta, granted entirely at Meta's discretion, not automatically. Below that bar, it's free to use commercially. The license also requires attribution ("Built with Llama") and Llama-branding in derivative products. One restriction did loosen over time: earlier versions banned using Llama's outputs to improve any other large language model outside the Llama family; starting with Llama 3.1, that's allowed as long as you attribute Llama in the resulting model and its documentation, which matters if distillation is part of your plan.

Weights are distributed through Hugging Face under the gated meta-llama org, requiring license acceptance before download, and through Meta's own developer channels. Serving support is Llama's clearest advantage: vLLM, Hugging Face TGI, Ollama, and llama.cpp all treat it as a first-class target, and it's natively available on every major managed platform, AWS Bedrock, Azure AI Foundry, Google Vertex, Together, Fireworks, Groq, and Databricks among them. The fine-tuning ecosystem is similarly the deepest of any open-weight family: more LoRA and QLoRA tooling has been built and tested against Llama first than against any other model, and it has by far the largest count of community fine-tunes published on Hugging Face, even in a year where its base model quality isn't the headline.

Where this fits at Numerata

If Llama's ecosystem maturity is what makes it the right choice for your team, P95 is where you fine-tune it on your own data, on infrastructure you control, private cloud or fully air-gapped, and NinetyFive serves it afterward, so you keep Llama's tooling advantage without giving up the latency or data control a managed API would cost you.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog