Gemma is Google DeepMind's open-weight model family, built from
the same research and technology as Gemini but released for anyone
to download and run, rather than served only through an API. Where
Qwen and Llama compete mostly on scale, Gemma has staked out a
narrower claim: the best capability per parameter at small sizes,
and the clearest path to actually running a model on a phone or
laptop rather than a GPU cluster. Here's who builds it, how it's
trained, and when that tradeoff is the right one.
Who builds Gemma
Google DeepMind builds Gemma as the open
counterpart to Gemini, not a separate research lineage; Google is
explicit that the two share underlying research and technology.
Gemma launched in February 2024 at 2B and 7B, followed by
Gemma 2 (9B and 27B, mid-2024), Gemma
3 (1B/4B/12B/27B, March 2025) as the first multimodal
generation, and Gemma 3n (June 2025) as a
dedicated on-device line for phones and browsers. Alongside the
general-purpose line, Google has shipped a steady stream of
specialized variants: CodeGemma, PaliGemma for vision-language
tasks, ShieldGemma for safety classification, EmbeddingGemma, and
others, each a smaller, purpose-built model rather than a
general-purpose checkpoint fine-tuned after the fact. Gemma 4
arrived in spring 2026, adding a mixture-of-experts size tier for
the first time.
A dense line through Gemma 3, an on-device-first branch with Gemma 3n, and Gemma 4's first mixture-of-experts tier.
How it's trained
Gemma 2 set the architectural pattern the dense
line still follows: it interleaves local sliding-window attention
with global attention every other layer to control the KV-cache
memory cost of longer context, on top of grouped-query attention.
Its 2B and 9B sizes were trained through knowledge
distillation from the 27B model as a teacher, rather than
pure next-token prediction on raw data, which is a meaningful part
of why the smaller sizes punch above their parameter count.
Gemma 3 (March 2025) carried that pattern forward
and added native multimodality through a SigLIP-based vision
encoder, increased the ratio of local-to-global attention layers
further to keep long-context memory in check, and moved to the
same 262K-entry SentencePiece tokenizer used in Gemini 2.0, tuned
to be more balanced across non-English text. Pretraining scale
grows with size, from roughly 2 trillion tokens at 1B up to about
14 trillion at 27B, with the larger sizes again trained via
distillation from bigger teacher models.
Gemma 3n (June 2025) is architecturally the
more interesting release: rather than just shrinking Gemma 3, it
introduces MatFormer, a nested "Matryoshka
Transformer" where a single trained model contains smaller,
independently valid sub-models inside it, and Per-Layer
Embeddings, which let per-layer embedding data stream in
from outside the model's main working memory rather than staying
resident the whole time. Together those let the E4B variant, 8
billion raw parameters, run with a memory footprint comparable to
a 4B model, which is where the "E" for "effective" in its name
comes from, and is what makes it viable to ship inside a phone app
rather than needing a server round-trip. Gemma 4 (2026)
ships as five variants built on that same lineage: E2B and E4B for
ultra-mobile use, a 12B model with a unified, encoder-free
architecture for multimodal input, a 26B-A4B mixture-of-experts
model that loads all 26B parameters but activates only 4B per
token, and a 31B dense model bridging server and local performance.
Every size natively handles text, image, and video, with audio
support specifically on the E2B, E4B, and 12B variants; context
runs 128K on the smaller models and up to 256K on the larger ones.
Gemma 4 also adds built-in multi-token prediction with dedicated
draft models for speculative decoding, and native system-prompt
support for more structured conversations. Post-training across
the family combines supervised
fine-tuning with reinforcement learning against a mix of reward
signals, human feedback, code-execution correctness, and math
ground-truth, aimed at reasoning, coding, multilingual quality, and
safety together rather than any one of them in isolation.
CAPABILITY PER PARAMETER
Google's own benchmarks put Gemma 3's 4B model competitive with the prior generation's 27B model, largely a product of the distillation-based training recipe carried across generations.
ON-DEVICE / MOBILE
Gemma 3n's MatFormer and per-layer embedding streaming are built specifically to run inside phone apps and browsers via MediaPipe and LiteRT, not just on a small server GPU.
SAFETY TOOLING
ShieldGemma ships as a dedicated safety-classifier model alongside the base line, giving teams a purpose-built moderation layer instead of having to build one from scratch.
When to use it
Gemma's size tiers run from the sub-1B range, used for
embeddings, function-calling classifiers, and other narrow tasks
like EmbeddingGemma and FunctionGemma, through 2B-4B and Gemma 4's
E2B/E4B variants for phone- and laptop-class inference, Gemma 3's
9B/12B or Gemma 4's unified 12B model as a strong
single-consumer-GPU general-purpose pick, up to Gemma 4's 31B
dense model or its 26B-A4B mixture-of-experts variant (26B total,
4B active per token) for the closest Gemma gets to
frontier-adjacent quality while remaining friendly to a single GPU
with quantization. There isn't a strong
case in 2026 commentary that Gemma is unconditionally the best
open-weight family at any given size, Llama still has the deepest
general tooling ecosystem, Qwen leads on coding and Chinese-language
tasks, Mistral is prized for raw serving speed, and Microsoft's
Phi line competes directly with Gemma at the smallest sizes. What
sets Gemma apart is less a benchmark score and more the deployment
story: if the target is a phone, browser, or single modest GPU,
Gemma's on-device tooling is more mature than any of those
alternatives offer today.
Reach for it specifically when you need real on-device inference
through MediaPipe or LiteRT, when you want a purpose-built safety
classifier alongside the base model rather than building one
yourself, or when single-GPU efficiency matters more than
squeezing out the last few points on a leaderboard.
Size tier tracks deployment fit, with the on-device tier being Gemma's most differentiated position.
How to use it
Licensing is worth checking per generation with Gemma, similar
to Mistral's history. Gemma 1 through 3, Gemma 3n, and the
specialized variants (CodeGemma, PaliGemma, ShieldGemma, and the
rest) all ship under Google's own Gemma Terms of Use
rather than a standard license like Apache 2.0, which incorporates
a separate Prohibited Use Policy by reference and requires anyone
redistributing the model to pass those same restrictions on to
their own users. Gemma 4 changed that: it ships
under a standalone Apache 2.0 license, a genuine
loosening of terms from the prior generations, so don't assume
Gemma 4's license from what an earlier Gemma release used.
Weights are available on Hugging Face under the
google org, on Kaggle, through Ollama's model
library, and via Google AI Studio, with community GGUF and
quantized variants widely available from groups like Unsloth. For
production serving, vLLM is the recommended path
with day-one support across NVIDIA and AMD GPUs as well as Google
Cloud TPUs; llama.cpp and Ollama
cover local and single-user workloads, though they trail vLLM
substantially on concurrent throughput. For genuinely on-device
deployment inside an app or browser, Google's own
MediaPipe LLM Inference API on LiteRT
is the first-party runtime built for exactly that, and it's the
piece none of Gemma's competitors currently match. Fine-tuning is
officially supported through both Keras and
Hugging Face's transformers/TRL stack, with Unsloth
as the common third-party option for LoRA and QLoRA at lower
memory cost.
Where this fits at Numerata
Gemma's efficiency-first design is a natural fit for
P95, where you fine-tune it on your
own data, on infrastructure you control, private cloud or fully
air-gapped, and NinetyFive serves it
afterward, so a model built to be capability-dense at small sizes
actually gets the low-latency serving that size makes possible,
instead of losing that advantage to a naive deployment.