An inference engine is the serving software that takes a
trained model's weights and runs them against live requests,
turning an input into a response. Training produces the weights.
The inference engine is the separate system that loads those
weights onto a GPU and executes the forward pass efficiently,
request after request, which is a different engineering problem
from training entirely.
Training and inference solve different problems
Training is about updating weights: running the model forward,
computing a loss, and backpropagating gradients through it, over
and over, until the weights converge. Inference never touches
gradients at all, it runs the forward pass exactly once per
request and returns the output. That sounds simpler, and
computationally it is, but inference has a constraint training
doesn't: a human or another system is waiting on the other end,
so latency is the whole game. A training job that takes an extra
ten minutes is a rounding error. An inference request that takes
an extra 200 milliseconds is a product problem.
Training loops for thousands of steps to shape weights. Inference runs one pass, and has to be fast every single time.
What the engine is actually doing
Loading a model's weights onto a GPU and calling forward on them
once would technically work, but it wouldn't scale past one request
at a time, and it would waste most of the GPU's capacity. An
inference engine exists to close that gap: it batches multiple
requests together so the GPU processes them in parallel instead of
one by one, manages the memory each request's running state
consumes so it can serve many concurrent users without running out
of VRAM, and schedules requests so a long one doesn't stall
everything behind it. None of that is visible in the model's
weights, it's entirely in the serving layer wrapped around them.
BATCHING
Groups requests that arrive close together so the GPU processes them as one batch instead of many small ones, which is dramatically more efficient use of the same hardware.
MEMORY MANAGEMENT
Each in-flight request holds state in GPU memory while it's being processed. The engine allocates and frees that memory precisely, so more concurrent requests fit without running out of VRAM.
SCHEDULING
Decides which requests run now versus queue, so one long request doesn't block a dozen short ones behind it, and latency stays predictable under real, uneven traffic.
Why the engine changes the latency you actually get
The same model weights can be served two, five, or ten times
faster depending entirely on the engine underneath them. Generic
serving stacks are usually built to be correct across almost any
model architecture, which means they carry overhead that a purpose
-built engine doesn't need to pay. Tricks like CUDA graphs, which
replay a pre-recorded sequence of GPU operations instead of
re-dispatching each one individually, can cut meaningful latency
off every request, but only if the engine is built to use them for
your specific model, rather than falling back to the general
case.
Where this fits at Numerata
NinetyFive is Numerata's inference
engine: it's what takes a model fine-tuned with
P95 and serves it at sub-50ms latency,
running inside your own trust boundary rather than a shared
third-party endpoint. It's the same serving path whether the model
is powering code completion, an internal chat tool, or an agent
workflow, just pointed at a different fine-tuned model each
time.