BLOG · EVALUATING

Should your coding assistant run on your own hardware?

Every engineering organization has now had this argument. One side points out that the assistant is the single most productive tool the team has adopted in a decade. The other side points out that it reads the entire repository and sends what it reads to someone else's servers. Both are correct, which is why the argument does not resolve on its own.

It resolves once you stop treating it as one question. A self-hosted AI coding assistant differs from a public one on exactly three axes: what crosses your network boundary, where the latency comes from, and how much of the system you can change. Everything else in the debate is either a restatement of one of those three or a cost question, which we have already worked through in the break-even math for self-hosting.

There are three arrangements, not two

Most comparisons set "public copilot" against "run your own" and skip the middle, which is where a lot of large firms actually sit. The middle tier matters because it changes the security answer without changing the architecture: the code still leaves, but a contract governs what happens to it afterward.

Three arrangements, and where source code crosses the network boundary Three stacked rows. In the consumer public assistant, the IDE sits inside your network and code crosses out to a shared vendor cloud, governed by default retention and training terms. In an enterprise tenancy, code still crosses out but into isolated tenancy under contracted zero-retention terms. In a self-hosted deployment, the IDE, the gateway, the inference server and the retrieval index all sit inside your network, and nothing crosses; model updates arrive as signed offline bundles. A · PUBLIC ASSISTANT, DEFAULT TIER YOUR NETWORK IDE + REPO CODE + CONTEXT SHARED VENDOR CLOUD default retention + training terms B · ENTERPRISE TENANCY YOUR NETWORK IDE + REPO CODE + CONTEXT ISOLATED TENANCY contracted zero retention, no training C · SELF-HOSTED YOUR NETWORK IDE + REPO GATEWAY MODEL SERVER PRIVATE RETRIEVAL INDEX nothing crosses; updates arrive as signed bundles
A and B share an architecture and differ by contract. C differs by architecture, which is why it is the only one that answers "prove it" with a packet capture.

Security: the four questions that actually differ

"Is it secure" is not answerable. Four narrower questions are, and they are the ones a review board will ask you. We wrote the full version of that questionnaire up separately in what your risk team will ask; these four are the subset specific to coding assistants.

What is actually in the payload. Teams reason about this as "the file I am editing," and it never is. A modern assistant assembles context: neighboring open tabs, imports resolved across the repository, recently viewed files, symbol definitions pulled by the language server, diagnostics, sometimes terminal output and test failures. That set is far broader than the cursor's file, and it is assembled by the extension rather than chosen by the developer. If you want to know what leaves, instrument the extension and read the requests. Most people are surprised once.

Retention, and for how long. Zero retention is a real and meaningful commitment. It is also frequently confused with "we do not train on it," which is a different promise. Ask for both in writing, ask what the abuse-monitoring exception retains, and ask how long that exception window is, because that window is the real retention period.

Which parties and which jurisdictions. A single completion may traverse a CDN, the vendor's inference provider, and a logging processor, each potentially in a different country. For a firm with data-residency obligations this is the question that decides the matter, and it is usually the one nobody asked until the audit.

Secrets and regulated data in context. Source code is not the only sensitive thing in a repository. Configuration files, fixtures, and test data routinely contain credentials, customer records, and account identifiers. Any assistant that reads broadly will eventually read those. This risk exists in all three arrangements; what differs is whether the leak is internal or external.

Dimension Public, default tier Enterprise tenancy Self-hosted
Code leaves your networkYesYesNo
Governed byTerms of serviceNegotiated contractYour own architecture
Trained on your codeSometimes, by defaultContractually noOnly if you train it
RetentionVendor-definedOften zero, with exceptionsWhatever you configure
Subprocessors and residencyVendor's choiceDisclosed, sometimes selectableNot applicable
Evidence available to auditorsVendor attestationsAttestations plus contractPacket capture, your own logs
Works air-gappedNoNoYes
Ops burden on youNoneMinimalReal, and ongoing

The last row is the honest cost of the last-but-one row. Being able to hand an auditor a week-long capture showing zero egress is worth a great deal in a regulated firm, and it is purchased with somebody's time.

Speed: what self-hosting does and does not fix

The performance claim for private copilots is usually stated too strongly. Self-hosting removes two specific components of latency and leaves the rest untouched.

It removes internet round trip, typically 40 to 120 milliseconds each way depending on where the nearest region sits, and replaces it with single-digit milliseconds on your own network. It also removes shared-tenant queueing, the variance that makes a public endpoint's median look excellent and its p99 look terrible at 9am in every timezone at once.

It does not make the model think faster. Prefill and decode take what they take, set by model size and your hardware, as covered in GPU sizing for inference. And if you under-provision, queueing comes back on your own hardware, where it is your problem rather than the vendor's.

Where the milliseconds go in one inline completion Two horizontal timelines for the same inline completion. Both begin with an identical client-side debounce and context assembly. The public path then adds internet round trip and shared queueing before prefill and decode. The self-hosted path on a local network has negligible round trip and a dedicated queue, so it finishes sooner even though the model work itself is identical. PUBLIC ENDPOINT OVER THE INTERNET DEBOUNCE CTX NET RTT QUEUE PREFILL DECODE SHOWN SELF-HOSTED ON YOUR OWN NETWORK DEBOUNCE CTX PREFILL DECODE SHOWN REMOVED BY SELF-HOSTING: INTERNET ROUND TRIP, SHARED-TENANT QUEUEING UNCHANGED: PREFILL AND DECODE, SET BY MODEL SIZE AND YOUR HARDWARE UNCHANGED: THE CLIENT-SIDE DEBOUNCE, WHICH IS OFTEN THE LARGEST SINGLE BLOCK
Two of the six blocks change. That is a real improvement in the tail, and it is not the order-of-magnitude story the category likes to tell.

Whether that improvement matters depends entirely on which mode of assistance you are talking about, because the three modes have completely different latency budgets.

Mode Useful budget What dominates Does self-hosting help
Inline completionUnder ~300ms to displayDebounce, round trip, a short decodeYes, materially, especially at p99
Chat in the editorUnder ~1s to first tokenPrefill over a long contextSomewhat: round trip is a small share
Agentic, multi-stepMinutes, measured end to endTotal tokens and number of tool callsBarely: throughput and model quality decide it

This is the single most useful thing to internalize about copilot performance. Inline completion is a latency product; agentic coding is a throughput and capability product. A private deployment serving a small fill-in-the-middle model on the local network is genuinely the better experience for the first. For the third, per-request latency is noise against a task that makes forty model calls, and the deciding factor is how good the model is, which is where the strongest public models still lead.

Which suggests the arrangement most large teams end up in: a self-hosted completion model for the constant, high-volume, code-adjacent traffic, and a deliberate decision about which model handles the long-horizon agentic work.

Customization: the part that compounds

Security is why private copilots get approved. Customization is why teams keep them. A public assistant is the same product for you as for everyone else, by design. A private one converges on your codebase.

Retrieval over what is actually yours. The most valuable context for a completion in a mature codebase is not on the public internet: it is the internal library that wraps your data access, the design doc explaining why the retry logic looks wrong but is not, the ticket where this exact edge case was decided. A private index over repositories, docs, and issue history is available to a self-hosted assistant and unavailable to a public one at any price.

Fine-tuning on your own idioms. Base models write generic code competently. They do not know that your team never uses a particular pattern, that internal calls go through your own client with a specific signature, or how your error types are meant to wrap. A modest fine-tune on your monorepo and merged pull requests moves acceptance rate more than a larger base model does, because most rejected suggestions are rejected for being unidiomatic rather than wrong. That is the work P95 exists to make routine, and the general shape is in what fine-tuning is.

Version pinning. Underrated, and the one that operations teams care about most once they have been bitten. On a public service the model changes when the vendor decides, and your prompts, your guardrails, and your evaluation baselines were all tuned against the old one. Self-hosted, the weights are a file. It changes when you change it, after your evaluation suite says the new one is better.

Routing. Once you own the gateway you can send completion traffic to a small fast model, chat to a mid-sized one, and reserve the expensive path for the requests that need it. This is usually where the economics of a private deployment actually work, because the overwhelming majority of requests do not need the biggest model.

Lever Public assistant Self-hosted
Private repo retrievalLimited to what the vendor indexesAny internal source you choose
Fine-tune on your codeRarely offeredYes, and the weights are yours
Pin a model versionVendor-controlledYes, it is a file
Route by request typeFixed product behaviorYours to define at the gateway
Custom context assemblyFixed by the extensionTunable, including redaction rules
Frontier capability todayBest availableBest open weights you can serve
Feature velocityContinuous, free to youWhatever you build or license

The last two rows are the honest counterweight, and they are the reason this post is not a recommendation to self-host everything.

Where public assistants still win

Four situations, stated plainly, where a public assistant is the right answer and a private copilot alternative is over-engineering.

The code is not sensitive. Plenty of internal tooling and greenfield product code carries no meaningful confidentiality risk. Treating it as though it does costs real money for no reduction in exposure.

The team is small. Per-seat pricing scales linearly and hardware does not, which means below some seat count the public option is simply cheaper. The arithmetic is in the cost post, and the threshold is higher than self-hosting advocates like to admit.

Nobody is asking. If no regulator, client contract, or internal policy constrains where code may be processed, the security argument is a preference rather than a requirement, and preferences do not justify an operations burden.

The work is frontier-shaped. Large refactors, multi-file agentic changes, and unfamiliar-language work lean on reasoning quality above everything else. The best public models are ahead there, and pretending otherwise leads to a private deployment your engineers quietly route around.

What a self-hosted copilot actually looks like

If you do go private, the architecture is not just "a model server somewhere." The pieces that make it usable, and reviewable, are the ones around the model.

Reference architecture for a self-hosted coding assistant Inside a single network boundary: IDE extensions send fill-in-the-middle and chat requests to a gateway that handles identity, policy, secret redaction and audit logging. The gateway routes completion traffic to a small fast model and chat or agent traffic to a larger model, both drawing on a retrieval index built over private repositories, docs and tickets. An evaluation and telemetry store feeds fine-tuning. Model updates enter as signed offline bundles; nothing exits. YOUR ENVIRONMENT IDE EXTENSION FIM + chat GATEWAY SSO + policy secret redaction audit log COMPLETION MODEL, SMALL latency path, most traffic CHAT / AGENT MODEL, LARGER throughput path RETRIEVAL INDEX · REPOS, DOCS, TICKETS, MERGED PRS EVAL + ACCEPTANCE TELEMETRY FINE-TUNE ON YOUR CODE FEEDBACK LOOP IN: SIGNED OFFLINE MODEL BUNDLES, ON YOUR SCHEDULE OUT: NOTHING
The model is the easy part. The gateway, the index, and the acceptance telemetry are what turn it into something engineers prefer to the public option.

Two components in that diagram get skipped and should not be. The gateway is what makes the deployment reviewable at all: identity from your existing directory, per-team policy, redaction before the prompt reaches the model, and an audit log you can export to your SIEM. And acceptance telemetry is the only honest measure of whether any of this is working. Log suggestion acceptance rate and edit distance after acceptance from day one. Without it you will be arguing about the assistant's value from anecdote, and the argument will not go your way.

How to decide, in one pass

Work down this list and stop at the first row that describes you. Most organizations know the answer by the third.

If this is trueThen
A regulator, client contract, or policy forbids code leaving your boundarySelf-host. The other axes are secondary
Your environment is air-gappedSelf-host; it is the only option that functions
You need the model to know your internal libraries and idiomsSelf-host, and budget for retrieval plus a fine-tune
Inline completion latency at p99 is the complaint you keep hearingSelf-host the completion model; leave the rest
Fewer than roughly 50 engineers, no compliance constraintEnterprise tenancy. Revisit at scale
The work is dominated by long agentic tasksKeep a strong public model in the mix, whatever else you run

Note that several of those rows point at a hybrid rather than a wholesale migration, and hybrid is the common landing point: private for the high-volume completion path and anything touching regulated code, public for the frontier-shaped work, with the gateway deciding which is which so the developer does not have to.

Five mistakes worth avoiding

Benchmarking on the wrong mode. A private deployment evaluated on long agentic tasks will lose to a frontier API, and one evaluated on inline completion will win. Whichever you measure, make sure it is the mode your engineers spend their day in.

Shipping without acceptance telemetry. If you cannot show that suggestion acceptance held or improved after the migration, the rollout will be judged on vibes and reversed.

Serving one model for everything. A single large model handling inline completion is slow and expensive at the same time. Split the paths.

Forgetting that context assembly is the product. Most of the perceived quality gap between assistants is context selection, not the model. If you self-host and keep a naive context window, engineers will notice a regression and correctly blame the migration.

Treating the security review as the finish line. Approval is when the work starts. Adoption is the outcome, and adoption is won on latency, idiom fit, and whether the thing knows about your internal libraries.

Where this fits at Numerata

Numerata is the infrastructure underneath the right-hand column of every table above. NinetyFive serves the models fast enough that a completion path on your own hardware beats a public endpoint on latency rather than merely matching it. P95 is where the fine-tune on your own repositories happens, and the resulting weights are yours and stay in your environment. Lupine keeps the cards from idling between the peaks. It installs inside your environment, private cloud, on-premises, or fully air-gapped, and the deployment includes the security review rather than leaving you to run it alone. If you want to know which of the three arrangements your situation actually calls for, that is the conversation we would rather have than a demo.

We have also covered the control-by-control version of the security column above, in an on-premises blueprint for preventing source code leakage: the nine paths code takes out of a network, what closes each one, and the evidence that proves it.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog