Every way source code leaves, and how to close each path.
AUGUST 11, 2026 · NUMERATA TEAM
Source code rarely leaks dramatically. There is usually no breach,
no attacker, and no incident report. A team adopts an AI tool, the
tool needs context to be useful, context means code, and code goes
wherever the tool's architecture sends it. By the time anyone asks
where, it has been going there for eight months.
This is a blueprint for running AI developer tools without that
happening: the paths code actually takes, the control that closes
each one, how far air-gapping goes, and the artifacts that let you
demonstrate the boundary rather than assert it. It assumes you have
already decided you want the tooling. If you are still weighing
whether to self-host at all, the
comparison of
the three arrangements is the better starting point, and this
post picks up where it leaves off.
Step 1: Inventory the paths code can take
Almost every leakage discussion starts and ends with the
completion request, which is the one path everybody already knows
about. It is not the one that surprises people eight months in.
The useful exercise is to list every component that can read
source, every destination each can reach, and whether a developer
would notice. That last column is what makes this worth writing
down, because the paths nobody notices are the ones that stay open.
Two paths a developer would notice. Seven they would not. Controls that only address path one leave a great deal open.
Write the equivalent table for your own stack before designing
any control, because a control you cannot map to a row is
decoration. Here is the general shape, with the control that closes
each path.
Path
What it carries
The control that closes it
Completion request
Open file, neighboring files, resolved symbols
Model served inside the boundary
Retrieval index
Embeddings of whole repositories, often all of them
Index inside the boundary, ACLs inherited from the repo
Prompt and reply logs
Verbatim code, held under different access rules
Content logging off by default, metadata only
Editor telemetry
File paths, repository names, sometimes snippets
Disabled by editor policy, blocked at egress
Crash reports
Stack traces containing source lines and paths
Reporting off, or routed to an internal collector
Extension installs
Arbitrary code with full workspace access
Pinned extension set from an internal registry
Agent tool calls
Anything the agent can read, anywhere it can reach
No session holds both file read and network reach
CI and build agents
Diffs, full files, secrets embedded in prompts
Same gateway, same policy, no direct API keys
Paste into a browser
Whatever the developer selected
A good internal option, plus browser DLP
Step 2: Put the model inside the boundary
Everything else in this post is detail work. This is the step
that changes the category of the guarantee.
When the model runs on hardware you control, the completion
request never reaches your network edge. There is no subprocessor
list to review, no retention setting to configure correctly, no
residency question to answer, and no contractual exception to read
twice. The property you want stops being something a vendor
promises and becomes something your topology enforces.
The distinction matters most in the room where it gets
questioned. A zero-retention agreement is a claim about what happens
to code that has already left; you can read the contract but you
cannot inspect the deletion. Absence of a network route is a claim
you can put a packet capture behind. Reviewers have learned the
difference, which is the subject of a
separate post on what
review boards ask.
The only crossing that exists is inbound, deliberate, and signed. Everything else is absent by policy rather than declined by agreement.
Step 3: Scope what the assistant may read
Moving the model inside closes the loud path. The retrieval index
is where the quiet one usually stays open, and it is worse than the
completion request in two ways: it holds whole repositories rather
than the file somebody happens to have open, and it holds them
persistently.
Two properties matter. The index lives inside the boundary, on
the same side as the model. And it enforces the same permissions as
the repositories it was built from, filtered at query time against
the identity of the person asking.
Miss the second and you have built something genuinely new: an
internal disclosure channel. The assistant will happily explain how
the pricing service authenticates to a developer who has no access
to that repository, because the index does not know it should
decline. No packet left your network. Something still leaked, and it
will be harder to detect than an egress problem because there is no
network artifact to find.
In practice that means the index stores the source repository and
path alongside each chunk, the gateway resolves group membership per
request from your identity provider, and the filter runs before
retrieval rather than after generation. Filtering the model's output
is not a control, because by then the content is in the context
window.
Step 4: Keep prompts and completions out of durable storage
A prompt in a coding assistant is source code. So a log
of prompts is a second copy of your repositories, usually with
weaker access controls than the repositories themselves, frequently
in a different system, and often with no retention policy because
nobody thought of it as code.
The default should be that content is not written down at all.
What you actually need for operating the service is metadata: which
user, which model version, token counts, latency, whether the
suggestion was accepted. That is enough to measure adoption, chase
regressions, and bill internally, and none of it is code.
When you do need content logging to debug a specific problem,
make it a deliberate, time-boxed switch: scoped to one user or one
repository, retained for days rather than indefinitely, stored in
the same tier as the source it duplicates, and logged as an event
itself so the enabling shows up in an audit trail.
Step 5: Close the side channels
This is the unglamorous half of the work, and the half that
determines whether the boundary is real.
Default-deny egress on the segment. Not a
firewall rule listing the destinations you want blocked, which is a
list you will always be behind on. An allowlist that starts empty
and stays empty for the assistant's own components.
DEFAULT-DENY EGRESS, KUBERNETES FORM
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: assistant-no-egress
namespace: ai-assistant
spec:
podSelector: {} # every pod in the namespace
policyTypes: [Egress]
egress:
# In-cluster only: the gateway, the index, the model server,# and DNS. No 0.0.0.0/0 rule exists, so nothing else resolves# or connects. Adding one later is a reviewable change.
- to:
- namespaceSelector:
matchLabels: { name: ai-assistant }
- to:
- namespaceSelector:
matchLabels: { name: kube-system }
ports:
- { protocol: UDP, port: 53 }
Pin the extension set. An IDE extension runs
with full workspace access, and marketplace installs are an
unreviewed supply chain into the machine holding your source. Serve
a fixed, reviewed set from an internal registry and disable
arbitrary installation by policy.
Turn off crash and usage reporting. In the
editor, in the language servers, and in the toolchain. Crash reports
are the most under-appreciated path on the list because they carry
stack traces with real source lines, and they fire precisely when
something unusual is happening.
Constrain agents so no session holds both
capabilities. An agent that can read the filesystem and
also fetch a URL, install a package, or call an external service is
an exfiltration channel by construction, with no malice required.
Scope filesystem access to the working repository, allowlist tools
explicitly rather than by exclusion, keep network-capable tools out
of sessions that read source, and require human approval for writes
outside the workspace.
Route CI through the same gateway. Build agents
that call a model directly with their own credentials are outside
every control you just built, and their prompts often contain both
diffs and secrets.
Deliver model updates as signed bundles. An
update mechanism that pulls over the network is a route in, and a
route in is a route. Bundles, verified by signature, applied when
you choose.
Accept that you cannot firewall the clipboard.
Browser DLP helps and is worth having. The durable control is that
the internal tool is good enough that pasting into a public chat is
not worth the trouble. Every deployment that treats this as purely
a policy problem gets shadow usage instead, which is the same leak
with less visibility.
Step 6: Prove the boundary holds
A diagram is a claim. A capture is evidence. The difference is
the whole of a security review, and assembling the evidence before
anyone asks is what turns a three-month review into a two-week
one.
THE ASSERTION WORTH RUNNING FOR A WEEK
# Capture at the segment uplink, excluding intra-segment traffic.# 10.42.0.0/16 is the assistant's own segment.
tcpdump -i uplink0 -nn -w egress.pcap \
'ip and not net 10.42.0.0/16'
# A week later, the number that matters. Anything above zero is# a finding: identify it, then decide if it belongs.
tcpdump -r egress.pcap -nn 2>/dev/null | wc -l
0
# Keep the pcap and this transcript. Reviewers ask for the# artifact, not the assurance that you ran it.
Run it over a period that includes a working week, a deploy, and
a model update, because those are the moments when something
unexpected tries to phone home. Then keep the output.
Three flavors of air-gapped, and what each costs
"Air-gapped" gets used for all three of these, which causes
trouble in review because they differ in what gets in, not in
whether code gets out. All three keep source inside. They cost
different amounts to operate.
Tier
Route to the internet
How updates arrive
What it costs you
Private cloud, isolated VPC
None from the AI subnet, default-deny at the edge
From your own mirror, over the network
Least. Cloud provider remains in scope for review
On-premises, connected
No internet route from the AI segment, internal reachable
Internal artifact mirror, pull on your schedule
Hardware and an internal mirror to maintain
Fully air-gapped
None, in either direction
Signed media, transferred deliberately
Slower updates, manual CVE tracking, real process
Pick the tier from the classification of the code, not from the
strictest thing available. A fully air-gapped deployment for a
repository that a contractor already has a laptop copy of is
expensive theater, and it competes for attention with controls that
would matter more. The
air-gapped training
post covers what changes in the pipeline when you do need the
third row.
The evidence pack
Six artifacts answer nearly every question a reviewer will ask
about a deployment like this. Assemble them once, keep them current,
and hand them over at the start of the review rather than producing
them one at a time under pressure.
Artifact
What it proves
Uplink packet capture, one week
No code left, demonstrated rather than asserted
Network policy manifests
The absence of egress is configuration, not luck
Log configuration plus sample records
What is recorded, and that content is not
Index ACL test results
A user cannot retrieve from a repo they cannot read
Agent tool allowlist
No session both reads source and reaches the network
Signed bundle manifest and verification transcript
The one inbound path is controlled and checkable
Five mistakes that reopen a closed boundary
Treating a contract as an architecture. A
zero-retention term and an absent network route are not the same
class of control. Buy the contract if it is the right tradeoff, but
do not record it as though the code stayed home.
Model inside, index outside. The most common
half-migration. The completion path looks clean on a diagram while
embeddings of every repository sit in a hosted vector database.
Forgetting that agents are egress. Controls
designed for autocomplete assume the tool only reads. An agent with
a shell and a network tool is a different threat model, and it is
usually adopted without a second review.
Logging prompts for quality and never revisiting
it. The switch gets flipped during a rollout to debug
acceptance rates, and two years later there is an unbounded store of
source code in an observability tool with its own access list.
Closing the sanctioned path harder than the unsanctioned
one. If the internal assistant is slow, poorly scoped, or
missing from the editor people actually use, they will use the
public one on a personal machine. That is the same leak, minus the
audit log.
Where this fits at Numerata
Numerata is how the middle column of the first table gets built.
NinetyFive serves the model inside the
boundary, fast enough that the sanctioned path is the one engineers
prefer, which is the control that outlasts policy.
P95 is where a model gets tuned on your
own repositories, with the resulting weights staying in your
environment. Lupine keeps the hardware
busy between peaks. It installs in your environment at whichever of
the three tiers above matches your code, and a
deployment includes the security
review as a named step, with the evidence pack assembled by us
rather than requested from us. There are no outbound connections by
construction, nothing phones home, and we hold no standing access to
your environment.