Should your coding assistant run on your own hardware?
AUGUST 10, 2026 · NUMERATA TEAM
Every engineering organization has now had this argument. One side
points out that the assistant is the single most productive tool the
team has adopted in a decade. The other side points out that it reads
the entire repository and sends what it reads to someone else's
servers. Both are correct, which is why the argument does not resolve
on its own.
It resolves once you stop treating it as one question. A
self-hosted AI coding assistant differs from a public
one on exactly three axes: what crosses your network boundary, where
the latency comes from, and how much of the system you can change.
Everything else in the debate is either a restatement of one of those
three or a cost question, which we have already worked through in
the break-even math for
self-hosting.
There are three arrangements, not two
Most comparisons set "public copilot" against "run your own" and
skip the middle, which is where a lot of large firms actually sit. The
middle tier matters because it changes the security answer without
changing the architecture: the code still leaves, but a contract
governs what happens to it afterward.
A and B share an architecture and differ by contract. C differs by architecture, which is why it is the only one that answers "prove it" with a packet capture.
Security: the four questions that actually differ
"Is it secure" is not answerable. Four narrower questions are, and
they are the ones a review board will ask you. We wrote the full
version of that questionnaire up separately in
what your risk team will
ask; these four are the subset specific to coding assistants.
What is actually in the payload. Teams reason about
this as "the file I am editing," and it never is. A modern assistant
assembles context: neighboring open tabs, imports resolved across the
repository, recently viewed files, symbol definitions pulled by the
language server, diagnostics, sometimes terminal output and test
failures. That set is far broader than the cursor's file, and it is
assembled by the extension rather than chosen by the developer. If you
want to know what leaves, instrument the extension and read the
requests. Most people are surprised once.
Retention, and for how long. Zero retention is a
real and meaningful commitment. It is also frequently confused with "we
do not train on it," which is a different promise. Ask for both in
writing, ask what the abuse-monitoring exception retains, and ask how
long that exception window is, because that window is the real
retention period.
Which parties and which jurisdictions. A single
completion may traverse a CDN, the vendor's inference provider, and a
logging processor, each potentially in a different country. For a firm
with data-residency obligations this is the question that decides the
matter, and it is usually the one nobody asked until the audit.
Secrets and regulated data in context. Source code
is not the only sensitive thing in a repository. Configuration files,
fixtures, and test data routinely contain credentials, customer
records, and account identifiers. Any assistant that reads broadly will
eventually read those. This risk exists in all three arrangements; what
differs is whether the leak is internal or external.
Dimension
Public, default tier
Enterprise tenancy
Self-hosted
Code leaves your network
Yes
Yes
No
Governed by
Terms of service
Negotiated contract
Your own architecture
Trained on your code
Sometimes, by default
Contractually no
Only if you train it
Retention
Vendor-defined
Often zero, with exceptions
Whatever you configure
Subprocessors and residency
Vendor's choice
Disclosed, sometimes selectable
Not applicable
Evidence available to auditors
Vendor attestations
Attestations plus contract
Packet capture, your own logs
Works air-gapped
No
No
Yes
Ops burden on you
None
Minimal
Real, and ongoing
The last row is the honest cost of the last-but-one row. Being able
to hand an auditor a week-long capture showing zero egress is worth a
great deal in a regulated firm, and it is purchased with somebody's
time.
Speed: what self-hosting does and does not fix
The performance claim for private copilots is usually stated too
strongly. Self-hosting removes two specific components of latency and
leaves the rest untouched.
It removes internet round trip, typically 40 to 120
milliseconds each way depending on where the nearest region sits, and
replaces it with single-digit milliseconds on your own network. It also
removes shared-tenant queueing, the variance that
makes a public endpoint's median look excellent and its p99 look
terrible at 9am in every timezone at once.
It does not make the model think faster. Prefill and decode take
what they take, set by model size and your hardware, as covered in
GPU sizing for
inference. And if you under-provision, queueing comes back on your
own hardware, where it is your problem rather than the vendor's.
Two of the six blocks change. That is a real improvement in the tail, and it is not the order-of-magnitude story the category likes to tell.
Whether that improvement matters depends entirely on which mode of
assistance you are talking about, because the three modes have
completely different latency budgets.
Mode
Useful budget
What dominates
Does self-hosting help
Inline completion
Under ~300ms to display
Debounce, round trip, a short decode
Yes, materially, especially at p99
Chat in the editor
Under ~1s to first token
Prefill over a long context
Somewhat: round trip is a small share
Agentic, multi-step
Minutes, measured end to end
Total tokens and number of tool calls
Barely: throughput and model quality decide it
This is the single most useful thing to internalize about copilot
performance. Inline completion is a latency product; agentic
coding is a throughput and capability product. A private
deployment serving a small fill-in-the-middle model on the local
network is genuinely the better experience for the first. For the
third, per-request latency is noise against a task that makes forty
model calls, and the deciding factor is how good the model is, which is
where the strongest public models still lead.
Which suggests the arrangement most large teams end up in: a
self-hosted completion model for the constant, high-volume, code-adjacent
traffic, and a deliberate decision about which model handles the
long-horizon agentic work.
Customization: the part that compounds
Security is why private copilots get approved. Customization is why
teams keep them. A public assistant is the same product for you as for
everyone else, by design. A private one converges on your codebase.
Retrieval over what is actually yours. The most
valuable context for a completion in a mature codebase is not on the
public internet: it is the internal library that wraps your data
access, the design doc explaining why the retry logic looks wrong but
is not, the ticket where this exact edge case was decided. A private
index over repositories, docs, and issue history is available to a
self-hosted assistant and unavailable to a public one at any price.
Fine-tuning on your own idioms. Base models write
generic code competently. They do not know that your team never uses a
particular pattern, that internal calls go through your own client with
a specific signature, or how your error types are meant to wrap. A
modest fine-tune on your monorepo and merged pull requests moves
acceptance rate more than a larger base model does, because most
rejected suggestions are rejected for being unidiomatic rather than
wrong. That is the work P95 exists to
make routine, and the general shape is in
what fine-tuning is.
Version pinning. Underrated, and the one that
operations teams care about most once they have been bitten. On a public
service the model changes when the vendor decides, and your prompts,
your guardrails, and your evaluation baselines were all tuned against
the old one. Self-hosted, the weights are a file. It changes when you
change it, after your evaluation suite says the new one is better.
Routing. Once you own the gateway you can send
completion traffic to a small fast model, chat to a mid-sized one, and
reserve the expensive path for the requests that need it. This is
usually where the economics of a private deployment actually work,
because the overwhelming majority of requests do not need the biggest
model.
Lever
Public assistant
Self-hosted
Private repo retrieval
Limited to what the vendor indexes
Any internal source you choose
Fine-tune on your code
Rarely offered
Yes, and the weights are yours
Pin a model version
Vendor-controlled
Yes, it is a file
Route by request type
Fixed product behavior
Yours to define at the gateway
Custom context assembly
Fixed by the extension
Tunable, including redaction rules
Frontier capability today
Best available
Best open weights you can serve
Feature velocity
Continuous, free to you
Whatever you build or license
The last two rows are the honest counterweight, and they are the
reason this post is not a recommendation to self-host everything.
Where public assistants still win
Four situations, stated plainly, where a public assistant is the
right answer and a private copilot alternative is over-engineering.
The code is not sensitive. Plenty of internal
tooling and greenfield product code carries no meaningful
confidentiality risk. Treating it as though it does costs real money
for no reduction in exposure.
The team is small. Per-seat pricing scales linearly
and hardware does not, which means below some seat count the public
option is simply cheaper. The arithmetic is in
the cost post, and the
threshold is higher than self-hosting advocates like to admit.
Nobody is asking. If no regulator, client contract,
or internal policy constrains where code may be processed, the security
argument is a preference rather than a requirement, and preferences do
not justify an operations burden.
The work is frontier-shaped. Large refactors,
multi-file agentic changes, and unfamiliar-language work lean on
reasoning quality above everything else. The best public models are
ahead there, and pretending otherwise leads to a private deployment
your engineers quietly route around.
What a self-hosted copilot actually looks like
If you do go private, the architecture is not just "a model server
somewhere." The pieces that make it usable, and reviewable, are the
ones around the model.
The model is the easy part. The gateway, the index, and the acceptance telemetry are what turn it into something engineers prefer to the public option.
Two components in that diagram get skipped and should not be. The
gateway is what makes the deployment reviewable at all:
identity from your existing directory, per-team policy, redaction before
the prompt reaches the model, and an audit log you can export to your
SIEM. And acceptance telemetry is the only honest
measure of whether any of this is working. Log suggestion acceptance
rate and edit distance after acceptance from day one. Without it you
will be arguing about the assistant's value from anecdote, and the
argument will not go your way.
How to decide, in one pass
Work down this list and stop at the first row that describes you.
Most organizations know the answer by the third.
If this is true
Then
A regulator, client contract, or policy forbids code leaving your boundary
Self-host. The other axes are secondary
Your environment is air-gapped
Self-host; it is the only option that functions
You need the model to know your internal libraries and idioms
Self-host, and budget for retrieval plus a fine-tune
Inline completion latency at p99 is the complaint you keep hearing
Self-host the completion model; leave the rest
Fewer than roughly 50 engineers, no compliance constraint
Enterprise tenancy. Revisit at scale
The work is dominated by long agentic tasks
Keep a strong public model in the mix, whatever else you run
Note that several of those rows point at a hybrid rather than a
wholesale migration, and hybrid is the common landing point: private for
the high-volume completion path and anything touching regulated code,
public for the frontier-shaped work, with the gateway deciding which is
which so the developer does not have to.
Five mistakes worth avoiding
Benchmarking on the wrong mode. A private deployment
evaluated on long agentic tasks will lose to a frontier API, and one
evaluated on inline completion will win. Whichever you measure, make
sure it is the mode your engineers spend their day in.
Shipping without acceptance telemetry. If you cannot
show that suggestion acceptance held or improved after the migration,
the rollout will be judged on vibes and reversed.
Serving one model for everything. A single large
model handling inline completion is slow and expensive at the same
time. Split the paths.
Forgetting that context assembly is the product.
Most of the perceived quality gap between assistants is context
selection, not the model. If you self-host and keep a naive context
window, engineers will notice a regression and correctly blame the
migration.
Treating the security review as the finish line.
Approval is when the work starts. Adoption is the outcome, and adoption
is won on latency, idiom fit, and whether the thing knows about your
internal libraries.
Where this fits at Numerata
Numerata is the infrastructure underneath the right-hand column of
every table above. NinetyFive serves the
models fast enough that a completion path on your own hardware beats a
public endpoint on latency rather than merely matching it.
P95 is where the fine-tune on your own
repositories happens, and the resulting weights are yours and stay in
your environment. Lupine keeps the cards
from idling between the peaks. It installs inside your environment,
private cloud, on-premises, or
fully air-gapped, and
the deployment includes the security review rather than leaving you to
run it alone. If you want to know which of the three arrangements your
situation actually calls for, that is the conversation we would rather
have than a demo.
We have also covered the control-by-control version of the security
column above, in
an on-premises
blueprint for preventing source code leakage: the nine paths code
takes out of a network, what closes each one, and the evidence that
proves it.