BLOG

What is an inference engine,
actually?

An inference engine is the serving software that takes a trained model's weights and runs them against live requests, turning an input into a response. Training produces the weights. The inference engine is the separate system that loads those weights onto a GPU and executes the forward pass efficiently, request after request, which is a different engineering problem from training entirely.

Training and inference solve different problems

Training is about updating weights: running the model forward, computing a loss, and backpropagating gradients through it, over and over, until the weights converge. Inference never touches gradients at all, it runs the forward pass exactly once per request and returns the output. That sounds simpler, and computationally it is, but inference has a constraint training doesn't: a human or another system is waiting on the other end, so latency is the whole game. A training job that takes an extra ten minutes is a rounding error. An inference request that takes an extra 200 milliseconds is a product problem.

Training loop versus inference request path Training is shown as a repeating loop of forward pass, loss, and backward pass feeding into updated weights over many iterations. Inference is shown as a single straight path from an incoming request through the model to a response, optimized for latency. TRAINING FORWARD LOSS BACKWARD REPEATS FOR THOUSANDS OF STEPS INFERENCE REQUEST ONE PASS RESPONSE <50MS
Training loops for thousands of steps to shape weights. Inference runs one pass, and has to be fast every single time.

What the engine is actually doing

Loading a model's weights onto a GPU and calling forward on them once would technically work, but it wouldn't scale past one request at a time, and it would waste most of the GPU's capacity. An inference engine exists to close that gap: it batches multiple requests together so the GPU processes them in parallel instead of one by one, manages the memory each request's running state consumes so it can serve many concurrent users without running out of VRAM, and schedules requests so a long one doesn't stall everything behind it. None of that is visible in the model's weights, it's entirely in the serving layer wrapped around them.

BATCHING

Groups requests that arrive close together so the GPU processes them as one batch instead of many small ones, which is dramatically more efficient use of the same hardware.

MEMORY MANAGEMENT

Each in-flight request holds state in GPU memory while it's being processed. The engine allocates and frees that memory precisely, so more concurrent requests fit without running out of VRAM.

SCHEDULING

Decides which requests run now versus queue, so one long request doesn't block a dozen short ones behind it, and latency stays predictable under real, uneven traffic.

Why the engine changes the latency you actually get

The same model weights can be served two, five, or ten times faster depending entirely on the engine underneath them. Generic serving stacks are usually built to be correct across almost any model architecture, which means they carry overhead that a purpose -built engine doesn't need to pay. Tricks like CUDA graphs, which replay a pre-recorded sequence of GPU operations instead of re-dispatching each one individually, can cut meaningful latency off every request, but only if the engine is built to use them for your specific model, rather than falling back to the general case.

Where this fits at Numerata

NinetyFive is Numerata's inference engine: it's what takes a model fine-tuned with P95 and serves it at sub-50ms latency, running inside your own trust boundary rather than a shared third-party endpoint. It's the same serving path whether the model is powering code completion, an internal chat tool, or an agent workflow, just pointed at a different fine-tuned model each time.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog