BLOG

What is a small language model,
actually?

A small language model (SLM) is a language model built to run with far fewer parameters than a frontier model, typically anywhere from a few hundred million up to around 10-15 billion, small enough to serve on a single GPU, or even a laptop or phone, with low latency. The tradeoff is breadth: a frontier model is trained to be good at almost anything. A small model is usually trained, or fine-tuned, to be very good at one thing.

Small relative to what

"Small" is relative, not a fixed cutoff. Frontier models from the largest labs run into the hundreds of billions of parameters, or use mixture-of-experts architectures that route between many billions more. Against that scale, anything in the single-digit or low double-digit billions counts as small. What matters more than the exact parameter count is what it implies: a small model fits in a fraction of the memory, serves far more requests per GPU, and responds in milliseconds instead of seconds.

Model size versus hardware footprint A large frontier model is shown as a tall stack of GPU blocks needed just to hold its weights, while a small language model is shown fitting on a single GPU block, at a small fraction of the size. FRONTIER MODEL GPU GPU GPU GPU MULTI-GPU CLUSTER, JUST TO SERVE SMALL LANGUAGE MODEL ONE GPU FITS ALONGSIDE OTHER WORKLOADS
Same task, wildly different footprint: a frontier model needs a cluster; a small model shares one GPU.

Why smaller can mean better, for a given task

A frontier model's size buys generality: it has to hold enough knowledge and reasoning ability to handle nearly any prompt a person might throw at it, in any domain, in any language. Most production use cases don't need that. A support bot answers questions about one product. A code-completion model writes in one codebase's style. An extraction pipeline pulls the same handful of fields out of the same kind of document, over and over. None of that requires general breadth, it requires depth on one narrow task, and a small model fine-tuned on exactly that task can match or beat a much larger general model at it, because none of its capacity is spent on things you'll never ask it to do.

LATENCY

Fewer parameters means fewer computations per token. A small model fine-tuned for your task can respond in the tens of milliseconds, where a frontier model call often takes a full second or more.

COST

Serving cost scales with model size. A small model on one GPU can handle far more concurrent requests than a frontier model spread across several, at a fraction of the per-request cost.

CONTROL

A small model is cheap enough to fine-tune and re-train often, and small enough to run entirely inside your own infrastructure, so the data that shapes it never has to leave your environment.

The catch: a small model needs a real task

The trade only pays off when the task is actually narrow enough to specialize for. A small base model with no fine-tuning is noticeably worse than a frontier model at almost everything, because it simply hasn't seen enough to be broadly capable. The gap closes, and often reverses, once you fine-tune it on data from the specific task you're deploying it for. That's the whole model: start from a small pretrained base, then narrow its remaining capacity toward one job with your own data.

Capability on a narrow task versus general breadth A chart showing a frontier model scoring high on general breadth but only moderately on one narrow task, while a small model fine-tuned for that task scores low on general breadth but highest on the narrow task itself. HIGH LOW FRONTIER: BREADTH FRONTIER: TASK SLM: BREADTH SLM: TASK
Fine-tuned for one task, a small model trades away breadth it wasn't using anyway.

Where this fits at Numerata

P95 is built for training and fine-tuning small models on your own data, so that specialization happens on infrastructure you control, private cloud or fully air-gapped. Once trained, those models serve through NinetyFive, Numerata's inference engine, which is what turns a small fine-tuned model's low parameter count into the sub-50ms latency it's actually capable of.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog