BLOG · NINETYFIVE ENGINEERING

Real-time inference doesn't
have minutes to spare.

torch.compile() promises real speedups from a single line of code, and it delivers. The cost is a warmup tax: the first request to a freshly compiled model can take minutes while PyTorch traces, optimizes, and generates kernels. For a real-time inference service where every millisecond of latency matters and deploys need to boot fast, that tradeoff doesn't work. So the NinetyFive team dug into what torch.compile() actually does under the hood, and found that nearly all of its speedup comes from one specific mechanism.

Where the speedup actually comes from

We benchmarked the time to generate one token on a 32-layer transformer across three configurations: plain eager execution, torch.compile() in reduce-overhead mode, and NVIDIA's CUDA Graphs API used directly, with no compiler involved at all.

Per-token latency across three execution modes Bar chart. Eager execution: 11.1 milliseconds. torch.compile in reduce-overhead mode: 8.1 milliseconds. Raw CUDA graphs with no compiler: 8.5 milliseconds, nearly matching torch.compile. EAGER 11.1 MS TORCH.COMPILE() 8.1 MS RAW CUDA GRAPHS 8.5 MS
Raw CUDA graphs get within 0.4ms of torch.compile(), with none of the compiler involved.

The speedup comes almost entirely from CUDA graphs eliminating kernel launch overhead. torch.compile() uses this API internally. Using it directly gets nearly the same steady-state performance, which raises an obvious question: what does the rest of the compiler actually buy you?

How CUDA graphs eliminate launch overhead

A CUDA graph records a sequence of GPU operations, kernel launches, memory copies, and so on, and replays them as a single unit. Normally, every kernel launch is a round trip: the CPU submits work, waits for the GPU to acknowledge, then submits the next piece. With hundreds of small kernels per forward pass, that overhead adds up fast. A CUDA graph collapses all of it into one command: replay this graph.

Kernel launches with and without a CUDA graph Top: eager execution, the CPU and GPU exchange a round trip for every kernel, hundreds of times per forward pass. Bottom: with a CUDA graph, the CPU issues a single replay command and the GPU runs the entire recorded kernel sequence internally. EAGER (NO CUDA GRAPH) CPU GPU ONE ROUND TRIP PER KERNEL LAUNCH, HUNDREDS OF TIMES PER FORWARD PASS WITH A CUDA GRAPH CPU graph.replay() KERNEL 1 → 2 → … → N ONE DISPATCH REPLAYS THE ENTIRE SEQUENCE
Eager execution pays a CPU-GPU round trip per kernel. A CUDA graph replays the whole sequence with one command.

CAPTURE ONCE, REPLAY REPEATEDLY

# Capture phase: record operations into a graph
graph = torch.cuda.CUDAGraph()
with torch.cuda.graph(graph):
    output = model(static_input)

# Replay phase: execute all operations with one command
static_input.copy_(new_input)
graph.replay()  # Runs the entire forward pass

The part torch.compile() throws away

Using CUDA graphs directly has another benefit that shows up immediately: first-request latency.

TORCH.COMPILE() FIRST REQUEST

4,152 ms spent tracing the model, running optimization passes, and generating kernels, on a single entry point. Multiply that across varying shapes and entry points and startup overhead reaches minutes.

RAW CUDA GRAPHS FIRST REQUEST

17 ms. No tracing, no optimization passes to run. Just a warmup call to let CUDA initialize its internal caches and allocate memory before anything gets captured.

CAN THIS GO LOWER

Yes. That 17ms warmup pass produces a perfectly valid result. Tools like torch.compile() just discard it instead of returning it.

So NinetyFive's serving layer does the obvious thing: return that warmup result to the caller, and defer graph capture to the next request that comes in.

Deferred graph capture across the first three requests Request one runs eager as a warmup pass and returns a valid result. Request two captures the CUDA graph while computing and still returns a valid result. Request three and every request after replays the captured graph at full speed. REQUEST 1 RUN EAGER (WARMUP) RETURN VALID RESULT REQUEST 2 CAPTURE GRAPH WHILE COMPUTING REQUEST 3+ REPLAY CAPTURED GRAPH FULL SPEED
Every request returns a valid result. Only the third and later requests get the fully replayed graph.

This brings warmup overhead down to something essentially imperceptible. It also happens to be less code: returning the warmup pass as a valid result is simpler than the alternative of discarding it and blocking until a graph is ready.

The variable-length input problem

CUDA graphs have one real complication. Capturing a graph records pointers to specific memory locations, and different tensor shapes get allocated to different memory. A graph captured for one shape can't be reused for another, which is a problem when input prompts can be any length.

torch.compile(dynamic=True) solves this by reasoning about shapes symbolically and generating kernels that work for any size, a much harder compiler problem, since it has to prove optimizations hold across every possible shape. NinetyFive's problem is narrower: only one dimension varies, the input sequence length. A naive fix would capture a separate graph, with its own buffers, for every sequence length seen, but that wastes memory quadratically. Instead, a single maximum-size buffer gets allocated once, and shorter sequences reuse slices of it.

Memory use with separate buffers versus one shared buffer Without sharing, three separate buffers of size 128, 256, and 512 are each allocated in full, totaling 896 units. With sharing, one 512-unit buffer is allocated once and sliced at 128 and 256 to serve all three sequence lengths, totaling 512 units. WITHOUT SHARING 128 256 512 896 UNITS ALLOCATED WITH SHARING 128 256 512 512 UNITS ALLOCATED (REUSED)
Three separate buffers cost 896 units. One shared buffer, sliced at each boundary, costs 512.

The one thing you can't know ahead of time is the largest input you'll actually see. Rather than pre-allocating for some assumed maximum, it's simpler to grow on demand: when a request arrives with a longer sequence than anything captured so far, invalidate the cached graphs and recapture with a larger shared buffer. In practice, a handful of long-prompt requests early in a session stabilize the buffer size for everything after.

Alternatives considered

PYTORCH'S BUILT-IN GRAPHED CALLABLES

torch.cuda.make_graphed_callables() does something similar, but its own warmup pass carries substantial overhead, so first-request latency stays high regardless.

TORCH.COMPILE(DYNAMIC=TRUE)

Solves variable shapes symbolically, generating kernels that work for any size. A much harder compiler problem than reusing slices of one buffer along a single dimension.

MEGAKERNELS

Hand-fusing everything into optimized CUDA can be extremely fast, but every new attention mechanism or sampling strategy means rewriting that CUDA by hand instead of iterating in Python.

Most of torch.compile()'s speedup comes from CUDA graphs eliminating kernel launch overhead. The rest of the compiler, operator fusion, custom Triton kernels, adds real startup time for comparatively small steady-state gains. By managing CUDA graphs directly and deferring capture instead of discarding the warmup pass, NinetyFive gets most of the performance with almost none of the startup penalty, on infrastructure you already control through Lupine and P95.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog