Real-time inference doesn't have minutes to spare.
JULY 29, 2026 · NINETYFIVE ENGINEERING TEAM
torch.compile() promises real speedups from a single line of code,
and it delivers. The cost is a warmup tax: the first request to a
freshly compiled model can take minutes while PyTorch traces,
optimizes, and generates kernels. For a real-time inference service
where every millisecond of latency matters and deploys need to boot
fast, that tradeoff doesn't work. So the NinetyFive team dug into
what torch.compile() actually does under the hood, and found that
nearly all of its speedup comes from one specific mechanism.
Where the speedup actually comes from
We benchmarked the time to generate one token on a 32-layer
transformer across three configurations: plain eager execution,
torch.compile() in reduce-overhead mode, and NVIDIA's CUDA Graphs API
used directly, with no compiler involved at all.
Raw CUDA graphs get within 0.4ms of torch.compile(), with none of the compiler involved.
The speedup comes almost entirely from CUDA graphs eliminating
kernel launch overhead. torch.compile() uses this API internally.
Using it directly gets nearly the same steady-state performance,
which raises an obvious question: what does the rest of the compiler
actually buy you?
How CUDA graphs eliminate launch overhead
A CUDA graph records a sequence of GPU operations, kernel
launches, memory copies, and so on, and replays them as a single
unit. Normally, every kernel launch is a round trip: the CPU submits
work, waits for the GPU to acknowledge, then submits the next piece.
With hundreds of small kernels per forward pass, that overhead adds
up fast. A CUDA graph collapses all of it into one command: replay
this graph.
Eager execution pays a CPU-GPU round trip per kernel. A CUDA graph replays the whole sequence with one command.
CAPTURE ONCE, REPLAY REPEATEDLY
# Capture phase: record operations into a graph
graph = torch.cuda.CUDAGraph()
with torch.cuda.graph(graph):
output = model(static_input)
# Replay phase: execute all operations with one command
static_input.copy_(new_input)
graph.replay() # Runs the entire forward pass
The part torch.compile() throws away
Using CUDA graphs directly has another benefit that shows up
immediately: first-request latency.
TORCH.COMPILE() FIRST REQUEST
4,152 ms spent tracing the model, running optimization passes, and generating kernels, on a single entry point. Multiply that across varying shapes and entry points and startup overhead reaches minutes.
RAW CUDA GRAPHS FIRST REQUEST
17 ms. No tracing, no optimization passes to run. Just a warmup call to let CUDA initialize its internal caches and allocate memory before anything gets captured.
CAN THIS GO LOWER
Yes. That 17ms warmup pass produces a perfectly valid result. Tools like torch.compile() just discard it instead of returning it.
So NinetyFive's serving layer does the
obvious thing: return that warmup result to the caller, and defer
graph capture to the next request that comes in.
Every request returns a valid result. Only the third and later requests get the fully replayed graph.
This brings warmup overhead down to something essentially
imperceptible. It also happens to be less code: returning the
warmup pass as a valid result is simpler than the alternative of
discarding it and blocking until a graph is ready.
The variable-length input problem
CUDA graphs have one real complication. Capturing a graph records
pointers to specific memory locations, and different tensor shapes
get allocated to different memory. A graph captured for one shape
can't be reused for another, which is a problem when input prompts
can be any length.
torch.compile(dynamic=True) solves this by reasoning about shapes
symbolically and generating kernels that work for any size, a much
harder compiler problem, since it has to prove optimizations hold
across every possible shape. NinetyFive's problem is narrower: only
one dimension varies, the input sequence length. A naive fix would
capture a separate graph, with its own buffers, for every sequence
length seen, but that wastes memory quadratically. Instead, a single
maximum-size buffer gets allocated once, and shorter sequences reuse
slices of it.
Three separate buffers cost 896 units. One shared buffer, sliced at each boundary, costs 512.
The one thing you can't know ahead of time is the largest input
you'll actually see. Rather than pre-allocating for some assumed
maximum, it's simpler to grow on demand: when a request arrives with
a longer sequence than anything captured so far, invalidate the
cached graphs and recapture with a larger shared buffer. In practice,
a handful of long-prompt requests early in a session stabilize the
buffer size for everything after.
Alternatives considered
PYTORCH'S BUILT-IN GRAPHED CALLABLES
torch.cuda.make_graphed_callables() does something similar, but its own warmup pass carries substantial overhead, so first-request latency stays high regardless.
TORCH.COMPILE(DYNAMIC=TRUE)
Solves variable shapes symbolically, generating kernels that work for any size. A much harder compiler problem than reusing slices of one buffer along a single dimension.
MEGAKERNELS
Hand-fusing everything into optimized CUDA can be extremely fast, but every new attention mechanism or sampling strategy means rewriting that CUDA by hand instead of iterating in Python.
Most of torch.compile()'s speedup comes
from CUDA graphs eliminating kernel launch overhead. The rest of the
compiler, operator fusion, custom Triton kernels, adds real startup
time for comparatively small steady-state gains. By managing CUDA
graphs directly and deferring capture instead of discarding the
warmup pass, NinetyFive
gets most of the performance with almost none of the startup
penalty, on infrastructure you already control through
Lupine
and P95.