BLOG · SLM SERIES

Gemma Explained

Gemma is Google DeepMind's open-weight model family, built from the same research and technology as Gemini but released for anyone to download and run, rather than served only through an API. Where Qwen and Llama compete mostly on scale, Gemma has staked out a narrower claim: the best capability per parameter at small sizes, and the clearest path to actually running a model on a phone or laptop rather than a GPU cluster. Here's who builds it, how it's trained, and when that tradeoff is the right one.

Who builds Gemma

Google DeepMind builds Gemma as the open counterpart to Gemini, not a separate research lineage; Google is explicit that the two share underlying research and technology. Gemma launched in February 2024 at 2B and 7B, followed by Gemma 2 (9B and 27B, mid-2024), Gemma 3 (1B/4B/12B/27B, March 2025) as the first multimodal generation, and Gemma 3n (June 2025) as a dedicated on-device line for phones and browsers. Alongside the general-purpose line, Google has shipped a steady stream of specialized variants: CodeGemma, PaliGemma for vision-language tasks, ShieldGemma for safety classification, EmbeddingGemma, and others, each a smaller, purpose-built model rather than a general-purpose checkpoint fine-tuned after the fact. Gemma 4 arrived in spring 2026, adding a mixture-of-experts size tier for the first time.

Gemma release timeline A timeline from Gemma 1 in February 2024 through Gemma 2 with knowledge distillation, Gemma 3 as the first multimodal generation, Gemma 3n for on-device use, to Gemma 4 in 2026 as the current generation. GEMMA 1 FEB 2024 GEMMA 2 MID-2024 GEMMA 3 / 3n 2025 GEMMA 4 2026
A dense line through Gemma 3, an on-device-first branch with Gemma 3n, and Gemma 4's first mixture-of-experts tier.

How it's trained

Gemma 2 set the architectural pattern the dense line still follows: it interleaves local sliding-window attention with global attention every other layer to control the KV-cache memory cost of longer context, on top of grouped-query attention. Its 2B and 9B sizes were trained through knowledge distillation from the 27B model as a teacher, rather than pure next-token prediction on raw data, which is a meaningful part of why the smaller sizes punch above their parameter count. Gemma 3 (March 2025) carried that pattern forward and added native multimodality through a SigLIP-based vision encoder, increased the ratio of local-to-global attention layers further to keep long-context memory in check, and moved to the same 262K-entry SentencePiece tokenizer used in Gemini 2.0, tuned to be more balanced across non-English text. Pretraining scale grows with size, from roughly 2 trillion tokens at 1B up to about 14 trillion at 27B, with the larger sizes again trained via distillation from bigger teacher models.

Gemma 3n (June 2025) is architecturally the more interesting release: rather than just shrinking Gemma 3, it introduces MatFormer, a nested "Matryoshka Transformer" where a single trained model contains smaller, independently valid sub-models inside it, and Per-Layer Embeddings, which let per-layer embedding data stream in from outside the model's main working memory rather than staying resident the whole time. Together those let the E4B variant, 8 billion raw parameters, run with a memory footprint comparable to a 4B model, which is where the "E" for "effective" in its name comes from, and is what makes it viable to ship inside a phone app rather than needing a server round-trip. Gemma 4 (2026) ships as five variants built on that same lineage: E2B and E4B for ultra-mobile use, a 12B model with a unified, encoder-free architecture for multimodal input, a 26B-A4B mixture-of-experts model that loads all 26B parameters but activates only 4B per token, and a 31B dense model bridging server and local performance. Every size natively handles text, image, and video, with audio support specifically on the E2B, E4B, and 12B variants; context runs 128K on the smaller models and up to 256K on the larger ones. Gemma 4 also adds built-in multi-token prediction with dedicated draft models for speculative decoding, and native system-prompt support for more structured conversations. Post-training across the family combines supervised fine-tuning with reinforcement learning against a mix of reward signals, human feedback, code-execution correctness, and math ground-truth, aimed at reasoning, coding, multilingual quality, and safety together rather than any one of them in isolation.

CAPABILITY PER PARAMETER

Google's own benchmarks put Gemma 3's 4B model competitive with the prior generation's 27B model, largely a product of the distillation-based training recipe carried across generations.

ON-DEVICE / MOBILE

Gemma 3n's MatFormer and per-layer embedding streaming are built specifically to run inside phone apps and browsers via MediaPipe and LiteRT, not just on a small server GPU.

SAFETY TOOLING

ShieldGemma ships as a dedicated safety-classifier model alongside the base line, giving teams a purpose-built moderation layer instead of having to build one from scratch.

When to use it

Gemma's size tiers run from the sub-1B range, used for embeddings, function-calling classifiers, and other narrow tasks like EmbeddingGemma and FunctionGemma, through 2B-4B and Gemma 4's E2B/E4B variants for phone- and laptop-class inference, Gemma 3's 9B/12B or Gemma 4's unified 12B model as a strong single-consumer-GPU general-purpose pick, up to Gemma 4's 31B dense model or its 26B-A4B mixture-of-experts variant (26B total, 4B active per token) for the closest Gemma gets to frontier-adjacent quality while remaining friendly to a single GPU with quantization. There isn't a strong case in 2026 commentary that Gemma is unconditionally the best open-weight family at any given size, Llama still has the deepest general tooling ecosystem, Qwen leads on coding and Chinese-language tasks, Mistral is prized for raw serving speed, and Microsoft's Phi line competes directly with Gemma at the smallest sizes. What sets Gemma apart is less a benchmark score and more the deployment story: if the target is a phone, browser, or single modest GPU, Gemma's on-device tooling is more mature than any of those alternatives offer today.

Reach for it specifically when you need real on-device inference through MediaPipe or LiteRT, when you want a purpose-built safety classifier alongside the base model rather than building one yourself, or when single-GPU efficiency matters more than squeezing out the last few points on a leaderboard.

Gemma size tiers mapped to deployment fit A chart mapping Gemma model sizes to deployment scenarios: sub-1 billion parameters for embeddings and narrow classifiers, Gemma 4's E2B and E4B on-device variants for phones and laptops, Gemma 3's 9B and 12B sizes or Gemma 4's unified 12B model for single consumer GPU serving, and Gemma 4's 31B dense model or its 26B-A4B mixture-of-experts variant for near-frontier quality still friendly to a single GPU. <1B EMBEDDINGS / CLASSIFIERS 2B-4B E2B/E4B, ON-DEVICE 9B-12B SINGLE-GPU SERVING 31B / 26B-A4B NEAR-FRONTIER
Size tier tracks deployment fit, with the on-device tier being Gemma's most differentiated position.

How to use it

Licensing is worth checking per generation with Gemma, similar to Mistral's history. Gemma 1 through 3, Gemma 3n, and the specialized variants (CodeGemma, PaliGemma, ShieldGemma, and the rest) all ship under Google's own Gemma Terms of Use rather than a standard license like Apache 2.0, which incorporates a separate Prohibited Use Policy by reference and requires anyone redistributing the model to pass those same restrictions on to their own users. Gemma 4 changed that: it ships under a standalone Apache 2.0 license, a genuine loosening of terms from the prior generations, so don't assume Gemma 4's license from what an earlier Gemma release used.

Weights are available on Hugging Face under the google org, on Kaggle, through Ollama's model library, and via Google AI Studio, with community GGUF and quantized variants widely available from groups like Unsloth. For production serving, vLLM is the recommended path with day-one support across NVIDIA and AMD GPUs as well as Google Cloud TPUs; llama.cpp and Ollama cover local and single-user workloads, though they trail vLLM substantially on concurrent throughput. For genuinely on-device deployment inside an app or browser, Google's own MediaPipe LLM Inference API on LiteRT is the first-party runtime built for exactly that, and it's the piece none of Gemma's competitors currently match. Fine-tuning is officially supported through both Keras and Hugging Face's transformers/TRL stack, with Unsloth as the common third-party option for LoRA and QLoRA at lower memory cost.

Where this fits at Numerata

Gemma's efficiency-first design is a natural fit for P95, where you fine-tune it on your own data, on infrastructure you control, private cloud or fully air-gapped, and NinetyFive serves it afterward, so a model built to be capability-dense at small sizes actually gets the low-latency serving that size makes possible, instead of losing that advantage to a naive deployment.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog