BLOG · EVALUATING

What it actually costs to run your own model

The honest answer is that running your own model is not automatically cheaper, and anyone who tells you otherwise is selling something. It is cheaper above a volume threshold, and the threshold moves depending on which API tier you would otherwise be paying for. Below that line, a hosted API is genuinely the cheaper option and you should use one.

What follows is the arithmetic, with every assumption stated so you can replace it with your own. The short version: against frontier-tier pricing, one GPU pays for itself somewhere around half a billion output tokens a month. Against the cheapest small hosted models, it may never pay for itself on token cost alone, and you should be buying it for control instead.

The two curves

An API bill is a straight line through the origin. You pay per token, so ten times the traffic costs ten times as much, and the line never flattens. Self-hosting is a staircase. You buy a GPU, and that GPU costs the same whether you push one request through it or saturate it. Cost per token falls as you fill the hardware, then jumps when you add the next card.

The interesting question is not which is cheaper. It is where the lines cross, and whether your volume is on the right side of that point.

API cost rises linearly with volume while self-hosted cost is a staircase A chart with monthly output tokens on the horizontal axis and monthly cost on the vertical. The API line rises in a straight line from the origin. The self-hosted line is flat, then steps up when a second GPU is added. The two lines cross at roughly 550 million tokens per month, marked as the break-even point. OUTPUT TOKENS PER MONTH COST API, PER TOKEN SELF-HOSTED + 2ND GPU BREAK-EVEN ~550M TOKENS/MO API CHEAPER HERE SELF-HOSTING CHEAPER HERE
The shape is what matters. Where the crossing point sits depends entirely on your assumptions, which is the rest of this post.

The assumptions

Every number below follows from these. Change one and the answer changes, which is exactly why you should run this with your own figures rather than trusting anyone's blog post, including this one.

InputAssumedWhy
Model8B class, fine-tuned, FP8The size most production classification, extraction, and scoring work lands on once it is tuned
Hardware1 x H100 80GB, reserved$1.80/hr reserved, roughly $1,314/month
Throughput3,000 output tok/secSustained with batching, not single-stream
Duty cycle30%Real traffic is peaky. Assuming 100% is the most common way this math gets faked
Effective capacity~2.3B output tok/month3,000/sec x 30% x 730 hrs
Operations0.25 FTE, ~$4,167/monthSomeone has to own upgrades, monitoring, and incidents
All-in per GPU~$5,481/monthHardware plus the quarter-engineer

Where the lines cross

With those inputs, one GPU costs $5,481 a month whether you use it or not. Divide that by the API price you would otherwise pay and you get the break-even volume.

Output tokens/month Self-hosted API at $10/M API at $2/M API at $0.50/M
10M$5,481$100$20$5
100M$5,481$1,000$200$50
550M$5,481$5,500$1,100$275
1B$5,481$10,000$2,000$500
2.3B (1 GPU full)$5,481$23,000$4,600$1,150
4.6B (2 GPUs)$6,795$46,000$9,200$2,300
23B (10 GPUs)$17,307$230,000$46,000$11,500

Three things fall out of that table, and only one of them is the one vendors like to quote.

Against frontier-tier pricing, self-hosting wins early. Break-even lands around 550 million output tokens a month, which a single busy internal application can reach. Past that the gap widens fast: at full utilization of one card you are paying $5,481 instead of $23,000.

Against mid-tier pricing, it wins late. At $2 per million you need to get past roughly 2.7 billion tokens a month, which is more than one card can serve, so you are buying the second GPU before the first one has paid off. The crossover is real but it is further out than most people assume.

Against the cheapest small hosted models, it may never win. One fully saturated H100 works out to about $2.34 per million output tokens all-in. If you can buy equivalent quality at $0.50, token cost alone will not justify the move, and you should be making the decision on control, latency, and data residency instead.

That last case is not a failure of the argument. It is the argument. The reason to run a model yourself is usually that the data cannot leave, or that the model needs to know things a general one does not. Cost is what makes that decision affordable at scale, not what triggers it.

What each bill actually contains

Both sides have costs that do not show up in the headline number, and both sides tend to leave out the other's.

What sits inside a self-hosted bill versus an API bill Two stacked columns. The self-hosted column contains GPU hours, operations time, storage and networking, and evaluation infrastructure. The API column contains per-token charges, retries and rate-limit handling, data egress, and exposure to price changes. SELF-HOSTED GPU HOURS OPS TIME STORAGE + NETWORK EVAL HARNESS HOSTED API PER-TOKEN CHARGES RETRIES + RATE LIMITS EGRESS + LOGGING PRICE-CHANGE RISK
The bottom three rows on each side are the ones missing from most comparisons.

The costs people forget

On the self-hosted side: idle time is the big one. A card at 10% duty cycle costs the same as one at 90%, so the cheapest thing you can do is consolidate workloads onto shared capacity rather than dedicating a GPU per team. After that: the engineer who owns upgrades, the evaluation infrastructure you need to prove a new model version is better before promoting it, and the storage for checkpoints, which grows faster than anyone plans for.

On the API side: retries against rate limits are billed tokens. So are the long system prompts you send on every single call to compensate for a model that was never trained on your domain, and those add up quietly. Then there is price-change exposure, which is not a line item but is a real risk when a workload becomes load-bearing and the price is set by somebody else.

When self-hosting does not pay

Be honest about these, because getting it wrong is expensive in both directions.

Spiky, low-volume traffic. If you are serving a few million tokens a month, an API is cheaper by an order of magnitude and you have better things to do with a quarter of an engineer.

You genuinely need frontier capability. If the task only works with the largest available model, a tuned 8B will not substitute for it, and the comparison is not like for like. Test that assumption though, because it is wrong more often than people expect on narrow, well-specified tasks.

Nobody owns it. Self-hosting with no named operator is how you end up with a model nobody has updated in a year serving production traffic. If you cannot name the person, the real cost is higher than the spreadsheet says.

How to run this yourself

Take your last month's API invoice and pull out the output token count, since that is what dominates. Divide your all-in monthly cost per GPU by your blended API price per million tokens. That gives you the break-even volume. Compare it against your current usage and your projected usage in twelve months, because infrastructure decisions are made against the second number.

If you land within roughly 2x of the crossing point in either direction, cost is not the deciding factor, and you should choose on data control, latency, and whether owning a tuned model on your own data is worth something to you.

Where this fits at Numerata

Duty cycle is the assumption doing the most work above, which is why Lupine scales to zero between jobs instead of holding reserved capacity, and why the same pool serves both training and inference rather than sitting idle in two places. NinetyFive is what determines the throughput number: better tokens per second per card moves your break-even point left. And our pricing is a flat yearly figure sized to what you run, so the staircase in that first diagram is the whole story. Your bill does not move when your traffic does.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog