BLOG · HOW-TO

Every way source code leaves,
and how to close each path.

Source code rarely leaks dramatically. There is usually no breach, no attacker, and no incident report. A team adopts an AI tool, the tool needs context to be useful, context means code, and code goes wherever the tool's architecture sends it. By the time anyone asks where, it has been going there for eight months.

This is a blueprint for running AI developer tools without that happening: the paths code actually takes, the control that closes each one, how far air-gapping goes, and the artifacts that let you demonstrate the boundary rather than assert it. It assumes you have already decided you want the tooling. If you are still weighing whether to self-host at all, the comparison of the three arrangements is the better starting point, and this post picks up where it leaves off.

Step 1: Inventory the paths code can take

Almost every leakage discussion starts and ends with the completion request, which is the one path everybody already knows about. It is not the one that surprises people eight months in.

The useful exercise is to list every component that can read source, every destination each can reach, and whether a developer would notice. That last column is what makes this worth writing down, because the paths nobody notices are the ones that stay open.

Nine paths source code can take out of a network Your repositories sit at the left. Nine outbound paths lead away from them. Two are visible to the developer: the completion request and pasting into a browser chat. Seven are quiet: the retrieval index, prompt and completion logs, editor telemetry, crash reports, extension installs, agent tool calls, and calls to a model API from CI jobs. YOUR REPOS VISIBLE TO THE DEVELOPER 1 · THE COMPLETION REQUEST current file, neighbors, symbols 2 · PASTE INTO A BROWSER whatever was on the clipboard QUIET 3 · RETRIEVAL INDEX embeddings of every repository 4 · PROMPT + REPLY LOGS a second copy, different ACLs 5 · EDITOR TELEMETRY file paths, repo names, snippets 6 · CRASH REPORTS stack traces with source lines 7 · AGENT TOOL CALLS reads a file, then fetches a URL 8 · CI JOBS + 9 · EXTENSIONS build-time API calls, marketplace
Two paths a developer would notice. Seven they would not. Controls that only address path one leave a great deal open.

Write the equivalent table for your own stack before designing any control, because a control you cannot map to a row is decoration. Here is the general shape, with the control that closes each path.

PathWhat it carriesThe control that closes it
Completion requestOpen file, neighboring files, resolved symbolsModel served inside the boundary
Retrieval indexEmbeddings of whole repositories, often all of themIndex inside the boundary, ACLs inherited from the repo
Prompt and reply logsVerbatim code, held under different access rulesContent logging off by default, metadata only
Editor telemetryFile paths, repository names, sometimes snippetsDisabled by editor policy, blocked at egress
Crash reportsStack traces containing source lines and pathsReporting off, or routed to an internal collector
Extension installsArbitrary code with full workspace accessPinned extension set from an internal registry
Agent tool callsAnything the agent can read, anywhere it can reachNo session holds both file read and network reach
CI and build agentsDiffs, full files, secrets embedded in promptsSame gateway, same policy, no direct API keys
Paste into a browserWhatever the developer selectedA good internal option, plus browser DLP

Step 2: Put the model inside the boundary

Everything else in this post is detail work. This is the step that changes the category of the guarantee.

When the model runs on hardware you control, the completion request never reaches your network edge. There is no subprocessor list to review, no retention setting to configure correctly, no residency question to answer, and no contractual exception to read twice. The property you want stops being something a vendor promises and becomes something your topology enforces.

The distinction matters most in the room where it gets questioned. A zero-retention agreement is a claim about what happens to code that has already left; you can read the contract but you cannot inspect the deletion. Absence of a network route is a claim you can put a packet capture behind. Reviewers have learned the difference, which is the subject of a separate post on what review boards ask.

The boundary, with the components inside it and the crossings that must not exist Inside the boundary: the IDE plugin, a gateway that authenticates and audits, the model server, the retrieval index, and an audit log. Outside: vendor telemetry, an extension marketplace, a hosted model API, and an observability service. Four crossing arrows are marked as blocked by default-deny egress. One inbound path is permitted: signed model bundles, verified and applied on your own schedule. YOUR NETWORK BOUNDARY IDE PLUGIN GATEWAY AUTH · AUDIT · QUOTA MODEL SERVER RETRIEVAL INDEX REPO ACLS AT QUERY TIME AUDIT LOG: METADATA, NOT CONTENT OUTSIDE VENDOR TELEMETRY EXTENSION MARKETPLACE HOSTED MODEL API OBSERVABILITY SAAS DEFAULT-DENY EGRESS: NO ROUTE OUT SIGNED MODEL BUNDLE VERIFIED, APPLIED ON YOUR SCHEDULE
The only crossing that exists is inbound, deliberate, and signed. Everything else is absent by policy rather than declined by agreement.

Step 3: Scope what the assistant may read

Moving the model inside closes the loud path. The retrieval index is where the quiet one usually stays open, and it is worse than the completion request in two ways: it holds whole repositories rather than the file somebody happens to have open, and it holds them persistently.

Two properties matter. The index lives inside the boundary, on the same side as the model. And it enforces the same permissions as the repositories it was built from, filtered at query time against the identity of the person asking.

Miss the second and you have built something genuinely new: an internal disclosure channel. The assistant will happily explain how the pricing service authenticates to a developer who has no access to that repository, because the index does not know it should decline. No packet left your network. Something still leaked, and it will be harder to detect than an egress problem because there is no network artifact to find.

In practice that means the index stores the source repository and path alongside each chunk, the gateway resolves group membership per request from your identity provider, and the filter runs before retrieval rather than after generation. Filtering the model's output is not a control, because by then the content is in the context window.

Step 4: Keep prompts and completions out of durable storage

A prompt in a coding assistant is source code. So a log of prompts is a second copy of your repositories, usually with weaker access controls than the repositories themselves, frequently in a different system, and often with no retention policy because nobody thought of it as code.

The default should be that content is not written down at all. What you actually need for operating the service is metadata: which user, which model version, token counts, latency, whether the suggestion was accepted. That is enough to measure adoption, chase regressions, and bill internally, and none of it is code.

When you do need content logging to debug a specific problem, make it a deliberate, time-boxed switch: scoped to one user or one repository, retained for days rather than indefinitely, stored in the same tier as the source it duplicates, and logged as an event itself so the enabling shows up in an audit trail.

Step 5: Close the side channels

This is the unglamorous half of the work, and the half that determines whether the boundary is real.

Default-deny egress on the segment. Not a firewall rule listing the destinations you want blocked, which is a list you will always be behind on. An allowlist that starts empty and stays empty for the assistant's own components.

DEFAULT-DENY EGRESS, KUBERNETES FORM

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: assistant-no-egress
  namespace: ai-assistant
spec:
  podSelector: {}            # every pod in the namespace
  policyTypes: [Egress]
  egress:
    # In-cluster only: the gateway, the index, the model server,
    # and DNS. No 0.0.0.0/0 rule exists, so nothing else resolves
    # or connects. Adding one later is a reviewable change.
    - to:
        - namespaceSelector:
            matchLabels: { name: ai-assistant }
    - to:
        - namespaceSelector:
            matchLabels: { name: kube-system }
      ports:
        - { protocol: UDP, port: 53 }

Pin the extension set. An IDE extension runs with full workspace access, and marketplace installs are an unreviewed supply chain into the machine holding your source. Serve a fixed, reviewed set from an internal registry and disable arbitrary installation by policy.

Turn off crash and usage reporting. In the editor, in the language servers, and in the toolchain. Crash reports are the most under-appreciated path on the list because they carry stack traces with real source lines, and they fire precisely when something unusual is happening.

Constrain agents so no session holds both capabilities. An agent that can read the filesystem and also fetch a URL, install a package, or call an external service is an exfiltration channel by construction, with no malice required. Scope filesystem access to the working repository, allowlist tools explicitly rather than by exclusion, keep network-capable tools out of sessions that read source, and require human approval for writes outside the workspace.

Route CI through the same gateway. Build agents that call a model directly with their own credentials are outside every control you just built, and their prompts often contain both diffs and secrets.

Deliver model updates as signed bundles. An update mechanism that pulls over the network is a route in, and a route in is a route. Bundles, verified by signature, applied when you choose.

Accept that you cannot firewall the clipboard. Browser DLP helps and is worth having. The durable control is that the internal tool is good enough that pasting into a public chat is not worth the trouble. Every deployment that treats this as purely a policy problem gets shadow usage instead, which is the same leak with less visibility.

Step 6: Prove the boundary holds

A diagram is a claim. A capture is evidence. The difference is the whole of a security review, and assembling the evidence before anyone asks is what turns a three-month review into a two-week one.

THE ASSERTION WORTH RUNNING FOR A WEEK

# Capture at the segment uplink, excluding intra-segment traffic.
# 10.42.0.0/16 is the assistant's own segment.
tcpdump -i uplink0 -nn -w egress.pcap \
    'ip and not net 10.42.0.0/16'

# A week later, the number that matters. Anything above zero is
# a finding: identify it, then decide if it belongs.
tcpdump -r egress.pcap -nn 2>/dev/null | wc -l
0

# Keep the pcap and this transcript. Reviewers ask for the
# artifact, not the assurance that you ran it.

Run it over a period that includes a working week, a deploy, and a model update, because those are the moments when something unexpected tries to phone home. Then keep the output.

Three flavors of air-gapped, and what each costs

"Air-gapped" gets used for all three of these, which causes trouble in review because they differ in what gets in, not in whether code gets out. All three keep source inside. They cost different amounts to operate.

TierRoute to the internetHow updates arriveWhat it costs you
Private cloud, isolated VPCNone from the AI subnet, default-deny at the edgeFrom your own mirror, over the networkLeast. Cloud provider remains in scope for review
On-premises, connectedNo internet route from the AI segment, internal reachableInternal artifact mirror, pull on your scheduleHardware and an internal mirror to maintain
Fully air-gappedNone, in either directionSigned media, transferred deliberatelySlower updates, manual CVE tracking, real process

Pick the tier from the classification of the code, not from the strictest thing available. A fully air-gapped deployment for a repository that a contractor already has a laptop copy of is expensive theater, and it competes for attention with controls that would matter more. The air-gapped training post covers what changes in the pipeline when you do need the third row.

The evidence pack

Six artifacts answer nearly every question a reviewer will ask about a deployment like this. Assemble them once, keep them current, and hand them over at the start of the review rather than producing them one at a time under pressure.

ArtifactWhat it proves
Uplink packet capture, one weekNo code left, demonstrated rather than asserted
Network policy manifestsThe absence of egress is configuration, not luck
Log configuration plus sample recordsWhat is recorded, and that content is not
Index ACL test resultsA user cannot retrieve from a repo they cannot read
Agent tool allowlistNo session both reads source and reaches the network
Signed bundle manifest and verification transcriptThe one inbound path is controlled and checkable

Five mistakes that reopen a closed boundary

Treating a contract as an architecture. A zero-retention term and an absent network route are not the same class of control. Buy the contract if it is the right tradeoff, but do not record it as though the code stayed home.

Model inside, index outside. The most common half-migration. The completion path looks clean on a diagram while embeddings of every repository sit in a hosted vector database.

Forgetting that agents are egress. Controls designed for autocomplete assume the tool only reads. An agent with a shell and a network tool is a different threat model, and it is usually adopted without a second review.

Logging prompts for quality and never revisiting it. The switch gets flipped during a rollout to debug acceptance rates, and two years later there is an unbounded store of source code in an observability tool with its own access list.

Closing the sanctioned path harder than the unsanctioned one. If the internal assistant is slow, poorly scoped, or missing from the editor people actually use, they will use the public one on a personal machine. That is the same leak, minus the audit log.

Where this fits at Numerata

Numerata is how the middle column of the first table gets built. NinetyFive serves the model inside the boundary, fast enough that the sanctioned path is the one engineers prefer, which is the control that outlasts policy. P95 is where a model gets tuned on your own repositories, with the resulting weights staying in your environment. Lupine keeps the hardware busy between peaks. It installs in your environment at whichever of the three tiers above matches your code, and a deployment includes the security review as a named step, with the evidence pack assembled by us rather than requested from us. There are no outbound connections by construction, nothing phones home, and we hold no standing access to your environment.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog