BLOG · EVALUATING

How many GPUs to serve a model at sub-50ms?

The answer comes from four numbers: how big the model is, what precision you run it at, how many requests are in flight at once, and what you actually mean by latency. Get those and the sizing falls out arithmetically. Most sizing exercises go wrong on the fourth one, so that is where we will start.

First, define the latency target

"Sub-50ms" is three different claims depending on what you measure, and conflating them is the most common mistake in capacity planning.

Time to first token (TTFT) is how long until the model emits anything. This is the prefill stage: it is compute bound and it scales with how long your prompt is.

Inter-token latency (ITL) is the gap between each subsequent token. This is the decode stage: it is memory bandwidth bound and essentially independent of prompt length.

Total round trip is TTFT plus ITL multiplied by the number of output tokens, plus network.

That last formula is the one that matters. A 200-token response at a very respectable 20ms per token takes four seconds, no matter how fast the first token arrived. So a sub-50ms total is only meaningful for short-output work: classification, extraction, scoring, ranking, signal generation. That covers a great deal of production finance work, and almost none of the chat demos people benchmark against. If you are generating paragraphs, stop quoting a total and start quoting TTFT and tokens per second, because those are the numbers you can actually hold yourself to.

The anatomy of a request: prefill then decode A timeline showing a request broken into a prefill block, which is compute bound and produces the first token, followed by a series of short decode steps, each memory bandwidth bound, one per output token. A short-output task finishes within 50 milliseconds while a long-output task continues for seconds. SHORT OUTPUT (CLASSIFY, EXTRACT, SCORE) PREFILL DONE < 50MS LONG OUTPUT (GENERATE PROSE) PREFILL ... x 200 TOKENS = SECONDS PREFILL = COMPUTE BOUND, SCALES WITH PROMPT LENGTH DECODE = BANDWIDTH BOUND, ONE PASS OVER THE WEIGHTS PER TOKEN
Same model, same hardware, same TTFT. Only the output length differs, and it is what decides whether 50ms is a plausible target.

Step 1: the memory floor

A GPU has to hold two things: the weights, and the KV cache for every sequence in flight. Weights are easy. Multiply parameters by bytes per parameter: an 8B model is 16GB at FP16, 8GB at FP8, 4GB at INT4.

KV cache is the one that surprises people, because it scales with context length and concurrency rather than with model size:

bytes/token = 2 x layers x kv_heads x head_dim x bytes_per_element

For an 8B-class model with 32 layers, 8 key-value heads under grouped-query attention, and head dimension 128, at FP16 that is 128 KiB per token. An 8,000-token sequence therefore needs about 1 GiB of cache entirely to itself. Thirty-two concurrent sequences at that context need 32 GiB, which is four times what the weights cost.

For a 70B model with 80 layers, the same arithmetic gives 320 KiB per token, so a single 8k sequence needs 2.5 GiB.

How an 80GB card is divided between weights and KV cache Two horizontal bars each representing 80 gigabytes of GPU memory. For an 8B model at FP8, weights take 8 gigabytes and about 70 gigabytes remain for KV cache. For a 70B model at FP8, weights take 70 gigabytes and only about 8 gigabytes remain, which is far fewer concurrent sequences. 8B @ FP8 ON ONE 80GB CARD 8GB ~70GB KV CACHE ~280 SEQUENCES @ 4K CONTEXT 70B @ FP8 ON ONE 80GB CARD 70GB WEIGHTS ~8GB ~3 SEQS THE MODEL FITTING IS NOT THE QUESTION. WHAT FITS ALONGSIDE IT IS.
A 70B model technically fits on one card. It just cannot serve anyone once it is there.

Step 2: single-stream speed

Decode is memory bandwidth bound, because the GPU reads the entire model out of memory to produce each token. So the ceiling is roughly:

tokens/sec = memory_bandwidth / model_size_in_bytes

An H100 SXM has about 3.35 TB/s of bandwidth. An 8B model at FP8 is 8GB, giving a theoretical 418 tokens per second single-stream. Real implementations reach 60 to 70% of theoretical, so call it 250 to 290 tokens per second, which is an ITL of roughly 3.5 to 4ms.

That is where a sub-50ms budget for short outputs comes from: a few milliseconds of prefill on a short prompt, then five or ten tokens at 4ms each, plus network. A 70B at FP8 is 70GB, giving about 48 tokens per second theoretical and closer to 30 in practice, so an ITL of around 30ms. The same ten-token output is now 300ms, and no amount of hardware changes that arithmetic. It is set by model size.

Step 3: batching, and the tradeoff nobody mentions

Batching is what makes GPUs economical, because reading the weights once serves every sequence in the batch. Throughput climbs almost linearly with batch size until you run out of memory bandwidth or cache.

The catch is that batching raises tail latency. A request that arrives just after a batch starts waits for the next scheduling slot. Continuous batching helps a great deal, but the tradeoff does not disappear: tuning for maximum throughput and tuning for p99 latency pull in opposite directions. Decide which one you are buying before you tune, and size against p99 rather than the average, because the average will always look fine.

Putting it together

ModelPrecisionWeights KV per 8k seqSingle-stream80GB cards
8BFP1616 GB1.0 GiB~130 tok/s1
8BFP88 GB0.5 GiB~260 tok/s1
32BFP832 GB1.5 GiB~70 tok/s1
70BFP870 GB2.5 GiB~30 tok/s2
70BINT435 GB2.5 GiB~55 tok/s1

Then size for concurrency and redundancy rather than for fit. Take your peak concurrent requests, divide by how many sequences fit alongside the weights, round up, and add one card for failover. For most tuned 8B workloads that lands at two cards: one to serve the traffic, one so a reboot is not an outage.

Four mistakes worth avoiding

Sizing on average load. Inference traffic is peaky, and market-hours traffic especially so. Size on p99 concurrency.

Forgetting the KV cache. Teams size for weights, deploy, and discover their concurrency ceiling is a fraction of what they planned, because cache is what actually ran out.

Benchmarking with short prompts. Prefill scales with prompt length. If production sends 4k-token prompts and you benchmarked on 200, your TTFT measurement is meaningless.

Assuming a bigger model is required. Sizing a 70B deployment for a task a tuned 8B handles is the most expensive mistake on this list, and the easiest to test for. Run the eval first.

Where this fits at Numerata

The throughput numbers above are what NinetyFive exists to improve: the same model on the same card, served faster, means fewer cards for the same traffic. Lupine is what stops those cards idling between market hours. And the sizing exercise in this post is exactly what the architecture review in a deployment produces, against your real prompt lengths and your real concurrency rather than illustrative ones.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog