A small language model (SLM) is a language model built to run
with far fewer parameters than a frontier model, typically anywhere
from a few hundred million up to around 10-15 billion, small enough
to serve on a single GPU, or even a laptop or phone, with low
latency. The tradeoff is breadth: a frontier model is trained to be
good at almost anything. A small model is usually trained, or
fine-tuned, to be very good at one thing.
Small relative to what
"Small" is relative, not a fixed cutoff. Frontier models from the
largest labs run into the hundreds of billions of parameters, or use
mixture-of-experts architectures that route between many billions
more. Against that scale, anything in the single-digit or low
double-digit billions counts as small. What matters more than the
exact parameter count is what it implies: a small model fits in a
fraction of the memory, serves far more requests per GPU, and
responds in milliseconds instead of seconds.
Same task, wildly different footprint: a frontier model needs a cluster; a small model shares one GPU.
Why smaller can mean better, for a given task
A frontier model's size buys generality: it has to hold enough
knowledge and reasoning ability to handle nearly any prompt a person
might throw at it, in any domain, in any language. Most production
use cases don't need that. A support bot answers questions about one
product. A code-completion model writes in one codebase's style. An
extraction pipeline pulls the same handful of fields out of the same
kind of document, over and over. None of that requires general
breadth, it requires depth on one narrow task, and a small model
fine-tuned on exactly that task can match or beat a much larger
general model at it, because none of its capacity is spent on
things you'll never ask it to do.
LATENCY
Fewer parameters means fewer computations per token. A small model fine-tuned for your task can respond in the tens of milliseconds, where a frontier model call often takes a full second or more.
COST
Serving cost scales with model size. A small model on one GPU can handle far more concurrent requests than a frontier model spread across several, at a fraction of the per-request cost.
CONTROL
A small model is cheap enough to fine-tune and re-train often, and small enough to run entirely inside your own infrastructure, so the data that shapes it never has to leave your environment.
The catch: a small model needs a real task
The trade only pays off when the task is actually narrow enough
to specialize for. A small base model with no fine-tuning is
noticeably worse than a frontier model at almost everything, because
it simply hasn't seen enough to be broadly capable. The gap closes,
and often reverses, once you fine-tune it on data from the specific
task you're deploying it for. That's the whole model: start from a
small pretrained base, then narrow its remaining capacity toward one
job with your own data.
Fine-tuned for one task, a small model trades away breadth it wasn't using anyway.
Where this fits at Numerata
P95 is built for training and
fine-tuning small models on your own data, so that specialization
happens on infrastructure you control, private cloud or fully
air-gapped. Once trained, those models serve through
NinetyFive, Numerata's inference
engine, which is what turns a small fine-tuned model's low
parameter count into the sub-50ms latency it's actually capable
of.