The answer comes from four numbers: how big the model is, what
precision you run it at, how many requests are in flight at once, and
what you actually mean by latency. Get those and the sizing falls out
arithmetically. Most sizing exercises go wrong on the fourth one, so
that is where we will start.
First, define the latency target
"Sub-50ms" is three different claims depending on what you measure,
and conflating them is the most common mistake in capacity
planning.
Time to first token (TTFT) is how long until the
model emits anything. This is the prefill stage: it is compute bound
and it scales with how long your prompt is.
Inter-token latency (ITL) is the gap between each
subsequent token. This is the decode stage: it is memory bandwidth
bound and essentially independent of prompt length.
Total round trip is TTFT plus ITL multiplied by the
number of output tokens, plus network.
That last formula is the one that matters. A 200-token response at a
very respectable 20ms per token takes four seconds, no matter how fast
the first token arrived. So a sub-50ms total is only meaningful
for short-output work: classification, extraction, scoring,
ranking, signal generation. That covers a great deal of production
finance work, and almost none of the chat demos people benchmark
against. If you are generating paragraphs, stop quoting a total and
start quoting TTFT and tokens per second, because those are the numbers
you can actually hold yourself to.
Same model, same hardware, same TTFT. Only the output length differs, and it is what decides whether 50ms is a plausible target.
Step 1: the memory floor
A GPU has to hold two things: the weights, and the KV cache for
every sequence in flight. Weights are easy. Multiply parameters by
bytes per parameter: an 8B model is 16GB at FP16, 8GB at FP8, 4GB at
INT4.
KV cache is the one that surprises people, because it scales with
context length and concurrency rather than with model size:
bytes/token = 2 x layers x kv_heads x head_dim x bytes_per_element
For an 8B-class model with 32 layers, 8 key-value heads under
grouped-query attention, and head dimension 128, at FP16 that is
128 KiB per token. An 8,000-token sequence therefore
needs about 1 GiB of cache entirely to itself. Thirty-two concurrent
sequences at that context need 32 GiB, which is four times what the
weights cost.
For a 70B model with 80 layers, the same arithmetic gives 320 KiB
per token, so a single 8k sequence needs 2.5 GiB.
A 70B model technically fits on one card. It just cannot serve anyone once it is there.
Step 2: single-stream speed
Decode is memory bandwidth bound, because the GPU reads the entire
model out of memory to produce each token. So the ceiling is roughly:
An H100 SXM has about 3.35 TB/s of bandwidth. An 8B model at FP8 is
8GB, giving a theoretical 418 tokens per second single-stream. Real
implementations reach 60 to 70% of theoretical, so call it 250 to 290
tokens per second, which is an ITL of roughly 3.5 to 4ms.
That is where a sub-50ms budget for short outputs comes from: a few
milliseconds of prefill on a short prompt, then five or ten tokens at
4ms each, plus network. A 70B at FP8 is 70GB, giving about 48 tokens
per second theoretical and closer to 30 in practice, so an ITL of
around 30ms. The same ten-token output is now 300ms, and no amount of
hardware changes that arithmetic. It is set by model size.
Step 3: batching, and the tradeoff nobody mentions
Batching is what makes GPUs economical, because reading the weights
once serves every sequence in the batch. Throughput climbs almost
linearly with batch size until you run out of memory bandwidth or
cache.
The catch is that batching raises tail latency. A request that
arrives just after a batch starts waits for the next scheduling slot.
Continuous batching helps a great deal, but the tradeoff does not
disappear: tuning for maximum throughput and tuning for
p99 latency pull in opposite directions. Decide which one you
are buying before you tune, and size against p99 rather than the
average, because the average will always look fine.
Putting it together
Model
Precision
Weights
KV per 8k seq
Single-stream
80GB cards
8B
FP16
16 GB
1.0 GiB
~130 tok/s
1
8B
FP8
8 GB
0.5 GiB
~260 tok/s
1
32B
FP8
32 GB
1.5 GiB
~70 tok/s
1
70B
FP8
70 GB
2.5 GiB
~30 tok/s
2
70B
INT4
35 GB
2.5 GiB
~55 tok/s
1
Then size for concurrency and redundancy rather than for fit. Take
your peak concurrent requests, divide by how many sequences fit
alongside the weights, round up, and add one card for failover.
For most tuned 8B workloads that lands at two cards: one to serve
the traffic, one so a reboot is not an outage.
Four mistakes worth avoiding
Sizing on average load. Inference traffic is
peaky, and market-hours traffic especially so. Size on p99
concurrency.
Forgetting the KV cache. Teams size for weights,
deploy, and discover their concurrency ceiling is a fraction of what
they planned, because cache is what actually ran out.
Benchmarking with short prompts. Prefill scales
with prompt length. If production sends 4k-token prompts and you
benchmarked on 200, your TTFT measurement is meaningless.
Assuming a bigger model is required. Sizing a 70B
deployment for a task a tuned 8B handles is the most expensive mistake
on this list, and the easiest to test for. Run the eval first.
Where this fits at Numerata
The throughput numbers above are what
NinetyFive exists to improve: the same
model on the same card, served faster, means fewer cards for the same
traffic. Lupine is what stops those cards
idling between market hours. And the sizing exercise in this post is
exactly what the architecture review in a
deployment produces, against your
real prompt lengths and your real concurrency rather than
illustrative ones.