The honest answer is that running your own model is not
automatically cheaper, and anyone who tells you otherwise is selling
something. It is cheaper above a volume threshold, and the threshold
moves depending on which API tier you would otherwise be paying for.
Below that line, a hosted API is genuinely the cheaper option and you
should use one.
What follows is the arithmetic, with every assumption stated so you
can replace it with your own. The short version: against frontier-tier
pricing, one GPU pays for itself somewhere around half a billion output
tokens a month. Against the cheapest small hosted models, it may never
pay for itself on token cost alone, and you should be buying it for
control instead.
The two curves
An API bill is a straight line through the origin. You pay per
token, so ten times the traffic costs ten times as much, and the line
never flattens. Self-hosting is a staircase. You buy a GPU, and that
GPU costs the same whether you push one request through it or saturate
it. Cost per token falls as you fill the hardware, then jumps when you
add the next card.
The interesting question is not which is cheaper. It is where the
lines cross, and whether your volume is on the right side of that
point.
The shape is what matters. Where the crossing point sits depends entirely on your assumptions, which is the rest of this post.
The assumptions
Every number below follows from these. Change one and the answer
changes, which is exactly why you should run this with your own
figures rather than trusting anyone's blog post, including this
one.
Input
Assumed
Why
Model
8B class, fine-tuned, FP8
The size most production classification, extraction, and scoring work lands on once it is tuned
Hardware
1 x H100 80GB, reserved
$1.80/hr reserved, roughly $1,314/month
Throughput
3,000 output tok/sec
Sustained with batching, not single-stream
Duty cycle
30%
Real traffic is peaky. Assuming 100% is the most common way this math gets faked
Effective capacity
~2.3B output tok/month
3,000/sec x 30% x 730 hrs
Operations
0.25 FTE, ~$4,167/month
Someone has to own upgrades, monitoring, and incidents
All-in per GPU
~$5,481/month
Hardware plus the quarter-engineer
Where the lines cross
With those inputs, one GPU costs $5,481 a month whether you use it
or not. Divide that by the API price you would otherwise pay and you
get the break-even volume.
Output tokens/month
Self-hosted
API at $10/M
API at $2/M
API at $0.50/M
10M
$5,481
$100
$20
$5
100M
$5,481
$1,000
$200
$50
550M
$5,481
$5,500
$1,100
$275
1B
$5,481
$10,000
$2,000
$500
2.3B (1 GPU full)
$5,481
$23,000
$4,600
$1,150
4.6B (2 GPUs)
$6,795
$46,000
$9,200
$2,300
23B (10 GPUs)
$17,307
$230,000
$46,000
$11,500
Three things fall out of that table, and only one of them is the
one vendors like to quote.
Against frontier-tier pricing, self-hosting wins early.
Break-even lands around 550 million output tokens a month, which a
single busy internal application can reach. Past that the gap widens
fast: at full utilization of one card you are paying $5,481 instead of
$23,000.
Against mid-tier pricing, it wins late. At $2 per
million you need to get past roughly 2.7 billion tokens a month, which
is more than one card can serve, so you are buying the second GPU
before the first one has paid off. The crossover is real but it is
further out than most people assume.
Against the cheapest small hosted models, it may never
win. One fully saturated H100 works out to about $2.34 per
million output tokens all-in. If you can buy equivalent quality at
$0.50, token cost alone will not justify the move, and you should be
making the decision on control, latency, and data residency
instead.
That last case is not a failure of the argument. It is the argument.
The reason to run a model yourself is usually that the data cannot
leave, or that the model needs to know things a general one does not.
Cost is what makes that decision affordable at scale, not what
triggers it.
What each bill actually contains
Both sides have costs that do not show up in the headline number,
and both sides tend to leave out the other's.
The bottom three rows on each side are the ones missing from most comparisons.
The costs people forget
On the self-hosted side: idle time is the big one.
A card at 10% duty cycle costs the same as one at 90%, so the cheapest
thing you can do is consolidate workloads onto shared capacity rather
than dedicating a GPU per team. After that: the engineer who owns
upgrades, the evaluation infrastructure you need to prove a new model
version is better before promoting it, and the storage for
checkpoints, which grows faster than anyone plans for.
On the API side: retries against rate limits are
billed tokens. So are the long system prompts you send on every single
call to compensate for a model that was never trained on your domain,
and those add up quietly. Then there is price-change exposure, which is
not a line item but is a real risk when a workload becomes load-bearing
and the price is set by somebody else.
When self-hosting does not pay
Be honest about these, because getting it wrong is expensive in both
directions.
Spiky, low-volume traffic. If you are serving a few
million tokens a month, an API is cheaper by an order of magnitude and
you have better things to do with a quarter of an engineer.
You genuinely need frontier capability. If the task
only works with the largest available model, a tuned 8B will not
substitute for it, and the comparison is not like for like. Test that
assumption though, because it is wrong more often than people expect
on narrow, well-specified tasks.
Nobody owns it. Self-hosting with no named operator
is how you end up with a model nobody has updated in a year serving
production traffic. If you cannot name the person, the real cost is
higher than the spreadsheet says.
How to run this yourself
Take your last month's API invoice and pull out the output token
count, since that is what dominates. Divide your all-in monthly cost
per GPU by your blended API price per million tokens. That gives you
the break-even volume. Compare it against your current usage and your
projected usage in twelve months, because infrastructure decisions are
made against the second number.
If you land within roughly 2x of the crossing point in either
direction, cost is not the deciding factor, and you should choose on
data control, latency, and whether owning a tuned model on your own
data is worth something to you.
Where this fits at Numerata
Duty cycle is the assumption doing the most work above, which is why
Lupine scales to zero between jobs
instead of holding reserved capacity, and why the same pool serves both
training and inference rather than sitting idle in two places.
NinetyFive is what determines the
throughput number: better tokens per second per card moves your
break-even point left. And our pricing is a flat yearly figure sized to
what you run, so the staircase in that first diagram is the whole
story. Your bill does not move when your traffic does.