BLOG · HOW-TO

How to deploy private AI

Deploying private AI means running a model's compute, fine-tuning, and inference entirely on infrastructure you control, private cloud or fully air-gapped, so your data and model weights never pass through a third-party API. In practice that comes down to five steps: decide how isolated you need to be, provision the compute, fine-tune the model on your own data, deploy an inference layer with a real latency budget, and verify there's no unintended network egress. Here's each step in detail.

Step 1: Decide how isolated you need to be

Every other decision in this guide depends on this one. Private cloud means a dedicated VPC or tenancy, isolated from other customers at the account level, but still connected to a provider's control plane and the internet for updates and telemetry. It's faster to stand up and covers most compliance requirements. Air-gapped means no network path in or out at all, verified rather than assumed, the requirement for classified data, heavily regulated environments, or anywhere a single unreviewed outbound connection is unacceptable on principle, not just on policy.

Choosing between private cloud and air-gapped deployment Your isolation requirement branches into two options. Private cloud: a dedicated VPC or tenancy, faster to stand up. Air-gapped: no network path in or out, required for classified or heavily regulated data. YOUR ISOLATION REQUIREMENT PRIVATE CLOUD Dedicated VPC or tenancy Faster to stand up AIR-GAPPED No network path in or out Classified / heavily regulated data
This choice determines every step that follows, so make it deliberately, not by default.

Step 2: Provision the compute

Once you know your isolation model, the compute itself has three practical paths: rack your own GPUs on-prem, rent isolated capacity from a private-cloud GPU provider, or use an on-demand pool that scales to zero when idle. None of these are inherently better, they just trade upfront cost against operational control. The one hard requirement is that whichever you pick has to sit inside the boundary you chose in step one. Renting a "private" GPU that still phones home for licensing defeats the point of an air-gapped requirement, and it's an easy detail to miss until an audit catches it.

Step 3: Fine-tune the model on your own data

Start from an open-weight base model, several current families ship under permissive licenses that allow commercial self-hosting, and fine-tune it on your own data rather than relying on prompting alone. Fine-tuning earns its cost specifically when you need consistent domain terminology or tone, a structured output format that has to hold up under edge cases, or latency and cost budgets a long few-shot prompt would blow on every request. For the mechanics of how that actually works, see what fine-tuning is and how it differs from prompting and RAG.

HEAD-ONLY

Fastest to train, ready the same day, and the right default when you just need to adapt a model's output format quickly.

LORA-MERGED

The best cost-to-accuracy balance for most classification and domain-adaptation tasks, and the default choice absent a specific reason to go further.

FULL FINE-TUNE

Highest accuracy ceiling, worth the extra cost specifically where wording and nuance carry the signal, not just keywords.

Step 4: Deploy an inference layer with a real latency budget

Serving is where a lot of private AI deployments quietly lose the advantage they built in steps one through three. Size the inference stack to your actual request volume and a specific latency target, not a generic default: an internal chat tool tolerates a few hundred milliseconds, but code completion or an agent loop calling the model dozens of times per task needs a serving layer built for real-time response, sub-50ms round trips rather than sub-second ones. Self-hosted inference frameworks like vLLM, SGLang, and TensorRT-LLM all handle the raw serving problem; the difference between a private deployment that feels fast and one that doesn't is almost always sizing and request batching strategy, not the framework choice itself.

Step 5: Verify there's no unintended network egress

This step gets skipped more often than any other, and it's the one an audit actually checks. A deployment can satisfy every requirement above and still have a network path nobody intended: a license check-in, an analytics call, a package that downloads model weights at runtime instead of using the copy already on disk. Test for these directly rather than assuming the absence of a complaint means the absence of a connection.

NO LICENSE CHECK-INS

Confirm the serving stack and any commercial tooling in the pipeline don't call home to validate a license at startup or on a schedule.

NO TELEMETRY

Disable and verify, don't just assume disabled, any analytics or crash-reporting SDKs bundled with your inference framework or its dependencies.

NO RUNTIME DOWNLOADS

Model weights, container images, and packages should be mirrored inside the boundary ahead of time, never pulled on demand from the open internet.

The five steps end to end, inside one isolation boundary A dashed boundary drawn from step one contains compute, fine-tuning, and serving from steps two through four, flowing to your application. Step five, verifying no egress, is checked at the boundary itself. STEP 1: YOUR ISOLATION BOUNDARY STEP 2 COMPUTE STEP 3 FINE-TUNE STEP 4 SERVE YOUR APP STEP 5: VERIFY NO EGRESS AT THIS LINE
All five steps end up as one boundary. Step five is what confirms it actually held.

Where this fits at Numerata

Lupine covers step two, on-demand GPUs that scale to zero, in our private cloud or fully air-gapped inside yours. P95 covers step three, fine-tuning on your own data with the same code running locally, in a sweep, or as a cloud job. NinetyFive covers step four, sub-50ms inference for whatever you're serving. Step one is a decision only you can make, and step five is a check worth running regardless of whose stack you deploy on.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog