BLOG · HOW-TO

How to deploy AI coding assistants in air-gapped and SOC 2 environments

Deploying an AI coding assistant in an air-gapped or SOC 2 environment is two problems wearing one label. Air-gapping decides where your source code is allowed to go. SOC 2 decides whether you can prove your controls actually work. A deployment can satisfy one and fail the other, so the first move is to stop treating them as the same requirement. From there the process is six steps: separate the two requirements, choose an assistant whose every component runs offline, move the model across the air gap without opening a path back, put a gateway in front for identity, redaction, and audit logging, map each control to a SOC 2 Trust Services Criterion, and prove zero egress with evidence you can hand an auditor. Here is each step in detail.

Air-gapped and SOC 2 answer two different questions

These two words get bundled into one procurement line so often that teams design for them as if they were a single control. They are not. Air-gapping is a network fact: there is no path in or out, verified rather than assumed. SOC 2 is an attestation: an independent auditor examined your controls and reported whether they were designed and operating effectively against the Trust Services Criteria. One is a property of your topology. The other is a property of your evidence.

The practical consequence is the thing most people get backwards. You can hold a clean SOC 2 report while running services that talk to the internet all day, because SOC 2 never asked you to disconnect, only to control and prove. And you can air-gap a deployment perfectly and still fail SOC 2, because pulling the plug says nothing about who can access the box, how changes are approved, or whether anyone is watching the logs. Air-gapping is a control you can choose. SOC 2 is proof that your controls, whichever you chose, operate.

Air-gapped versus SOC 2, as two different questions Two columns. The air-gapped column governs where data may go: its mechanism is a network boundary, it is satisfied by having no path in or out, and it is proven by a packet capture and deny-all firewall rules. The SOC 2 column governs whether controls work: its mechanism is documented controls plus evidence, it is satisfied by access management, change management and monitoring, and it is proven by an independent auditor's Type II report. A full-width strip below states that neither implies the other. AIR-GAPPED GOVERNS where data may go MECHANISM a network boundary SATISFIED BY no path in or out PROVEN BY packet capture, deny-all rules a property of your topology SOC 2 GOVERNS whether controls work MECHANISM controls + evidence SATISFIED BY access, change mgmt, monitoring PROVEN BY a Type II report a property of your evidence NEITHER IMPLIES THE OTHER air-gapping is a control you choose · SOC 2 is proof your controls operate
Design for both as one thing and you will over-build the network and under-build the evidence, or the reverse. They are separate requirements that happen to reinforce each other.

The good news is that they do reinforce each other. An air-gapped coding assistant makes the hardest SOC 2 questions easy to answer, because "does source code leave our boundary" stops being a policy argument and becomes a packet capture. The rest of this guide builds the deployment so that isolation and evidence come out of the same architecture rather than being bolted on twice.

Step 1: Separate the air-gap requirement from the SOC 2 requirement

Before any hardware or model choice, write down which of the two you are actually on the hook for, because the answer is often not both. A regulator or a customer contract may demand air-gapping for a specific class of code. A different customer may only demand a SOC 2 Type II report and never mention the network at all. Scope them independently and you avoid the two expensive mistakes: air-gapping an environment nobody required isolated, and assuming an isolated environment is automatically audit-ready when it has no access reviews behind it.

The distinction also decides who owns the work. Air-gapping is an infrastructure decision, mostly settled once. SOC 2 is a continuous program: controls that operate every day and generate evidence every day, sampled by an auditor over a review period. If you are pursuing both, treat the air gap as one control inside the SOC 2 program, not as a substitute for it. We walk the isolation choice itself in more depth in how to deploy private AI; this post assumes you have landed on air-gapped and focuses on doing it in a way an auditor can sign.

Step 2: Choose an assistant whose every component runs offline

An AI coding assistant is not one program. It is at least four, and each has to run with no outbound calls for the deployment to be genuinely air-gapped. The model weights and the inference server are the parts everyone checks. The two that quietly break isolation are the IDE extension and the retrieval index.

THE MODEL + SERVER

An open-weight model served by a self-hosted stack. The check: it loads weights from local disk, never fetches them at runtime, and its telemetry and update channels are off.

THE IDE EXTENSION

The usual failure point. Many extensions still send telemetry, license check-ins, or feature flags even in enterprise mode. You need one that points at your endpoint with no mandatory calls home.

THE RETRIEVAL INDEX

The context that makes the assistant good, built over your repos, docs, and tickets, lives inside the boundary. No hosted vector database, no embeddings API, no external reranker.

The IDE extension deserves the paranoia. A modern assistant does not send "the file I am editing." It assembles context: neighboring open tabs, imports resolved across the repository, symbol definitions from the language server, sometimes terminal output and failing test names. That whole bundle is what would leave if the extension had a path out, which in an air-gapped deployment it must not. Instrument the extension and read its requests before you trust the marketing; the full inventory of where code escapes a network is in every way source code leaves, and how to close each path.

Step 3: Move the model across the air gap without opening a path back

An air-gapped model still has to get in somehow, and "somehow" is exactly where isolation is usually lost. The wrong way is a temporary firewall exception to pull weights or a package during install, which is a path out that existed for an afternoon and will not appear in any diagram. The right way is a one-way transfer: stage everything on a connected machine, scan and sign it, carry the signed bundle across a controlled transfer, and import it inside the boundary with signature verification. Nothing about the import creates a return path.

Moving a model across the air gap as a one-way, signed transfer On the connected side, an open-weight model and its dependencies are staged, then scanned and signed. The signed bundle crosses a one-way transfer at the air gap. Inside the boundary the signature is verified and the model is served. There is no arrow back across the gap, and each crossing is recorded as a change ticket with approval. CONNECTED STAGING OPEN-WEIGHT MODEL + DEPS SCAN + SIGN SIGNED BUNDLE AIR GAP ONE WAY no return path YOUR BOUNDARY VERIFY SIGNATURE SERVE FROM LOCAL DISK each crossing = one change ticket, approved
The same one-way crossing that keeps the environment isolated is also your change-management evidence: a signature and an approval per model update, which is exactly what a SOC 2 auditor samples.

Notice what this step buys you twice. Operationally, it is how the model gets in without ever letting anything out. For the audit, the signature and the approval ticket on every crossing are the artifact that answers "how do you control changes to production models," which is a control an internet-connected deployment has to manufacture separately. The air gap forced you to build change management as a side effect. The full air-gapped workflow, including how updates keep arriving on your schedule, is covered in air-gapped LLM training.

Step 4: Put a gateway in front for identity, redaction, and audit logging

A model server on its own is not auditable. It answers whoever can reach it, redacts nothing, and remembers nothing. The component that turns a private model into a reviewable system is a gateway that every request passes through, and it is the single most load-bearing part of a SOC 2 deployment. It is where identity, policy, redaction, and the audit trail all live.

Reference architecture: a gateway as the control and evidence plane Inside one air-gapped boundary, IDE extensions send completion and chat requests to a gateway that handles single sign-on and role-based access, secret and PII redaction, and immutable audit logging. The gateway routes to a small completion model and a larger chat or agent model, both drawing on a private retrieval index over repos, docs and tickets. The audit log flows to your SIEM. Signed model bundles enter from step three; nothing exits. YOUR AIR-GAPPED BOUNDARY IDE EXTENSION completion + chat GATEWAY SSO + RBAC secret redaction audit log SOC 2 control plane COMPLETION MODEL, SMALL latency path, most traffic CHAT / AGENT MODEL, LARGER throughput path PRIVATE RETRIEVAL INDEX · REPOS, DOCS, TICKETS AUDIT LOG → YOUR SIEM ACCESS REVIEWS + EVIDENCE IN: SIGNED MODEL BUNDLES (STEP 3) OUT: NOTHING
The model is the easy part. The gateway is what an auditor actually interviews: it is where you show identity, redaction, and a log, and it is the one place a control failure would show up.

Three of the gateway's jobs are worth stating plainly, because each maps to a control an auditor will test. Identity and access: requests authenticate against your existing directory and are authorized per team, so "who could use the assistant, and on which repositories" has an answer you can export. Redaction: secrets and regulated fields are stripped before the prompt reaches the model, because configuration files and fixtures routinely carry credentials that no assistant should ingest. Audit logging: an immutable record of requests, shipped to your SIEM, which is both your monitoring evidence and, because it now holds source code, a confidential store in its own right.

That last point is the subtle one. The moment you log prompts and completions, the log contains code and possibly secrets, so the log store inherits the same access controls, retention limits, and encryption as production. An audit trail is a SOC 2 asset for monitoring and a SOC 2 liability for confidentiality at the same time, and the deployments that fail here are the ones that treated logging as free. This is the same control surface a risk team probes in a vendor review, which we lay out in what your risk team will ask.

Step 5: Map each control to a SOC 2 Trust Services Criterion

SOC 2 is built on five Trust Services Criteria: Security, Availability, Processing Integrity, Confidentiality, and Privacy. Security, known as the Common Criteria, is mandatory in every report. For a coding assistant, Confidentiality almost always applies because the prompts are source code, and Availability usually does because developers block on the tool. Processing Integrity and Privacy come into scope when the work calls for them. An audit does not grade your architecture; it grades whether each control you claim is real and operating. So the deliverable for this step is a table that ties every control in the deployment to a criterion and, critically, to the artifact that proves it.

Control in the deployment SOC 2 criterion Evidence you can produce
SSO and role-based access at the gatewayLogical access (CC6)Directory group membership, periodic access-review exports
Secret and regulated-data redaction before the modelConfidentiality (C1)Redaction rules in version control, sampled prompt logs
Immutable audit log shipped to your SIEMMonitoring (CC7)SIEM records, alert configuration, retention policy
Signed, reviewed model bundles imported on a scheduleChange management (CC8)Bundle signatures, change tickets, approval records
No outbound network path from the serving subnetConfidentiality (C1)Firewall deny-all rules, dated packet captures
Capacity and failover for the serving layerAvailability (A1)Uptime monitoring, load-test results, runbooks
Evaluation gate before a new model serves trafficProcessing integrity (PI1)Eval-suite results, sign-off before promotion

Read the table the way an auditor does, right column first. A control with no producible evidence is not a control; it is an intention. The reason an air-gapped deployment is comfortable to audit is that its strongest evidence, the deny-all rule and the capture showing nothing leaves, is objective and repeatable rather than a matter of trust. You are not asking the auditor to believe code stays inside. You are showing them the wall.

Step 6: Prove zero egress and keep the evidence auditors ask for

This is the step teams skip, and it is the one both the air-gap claim and the SOC 2 report actually rest on. Absence of a complaint is not absence of a connection. A deployment can satisfy every step above and still have a path nobody intended: a dependency that phones home on update, an analytics SDK bundled three layers deep, a container that pulls a base image at restart. Prove the boundary directly, and prove it more than once.

DENY BY DEFAULT

The serving subnet blocks outbound traffic as the default rule, not as an afterthought. Anything that needs to reach out has to be an explicit, reviewed exception, of which the target is zero.

CAPTURE AND KEEP

A packet capture over a real working session, showing no egress, dated and stored. This is the artifact that converts "we are air-gapped" from an assertion into evidence a Type II audit can sample.

RE-VERIFY ON CHANGE

Egress is not proven once at go-live. A model update, a dependency bump, or a config change can reopen a path, so the check runs again on every change and the result is filed.

The word that matters in SOC 2 Type II is period. Type I says your controls were designed correctly on the day the auditor looked. Type II says they operated correctly across a review window, often three to twelve months, sampled without warning. That is why the evidence in step five has to be standing rather than staged: access reviews that happen on a cadence, captures that get re-taken on change, tickets that exist because a process created them and not because someone assembled a folder the week before the audit. Isolation removes a class of risk in one architectural stroke; the report is earned by keeping the evidence flowing.

Five mistakes that fail the audit

Treating air-gapping as if it were SOC 2, or the reverse. The isolated environment with no access reviews and the compliant environment with an unexamined outbound path are the same mistake from opposite ends. Scope both, in step one.

Logging without putting the log in scope. Prompt and completion logs contain source code and secrets. If the log store does not inherit production access controls and retention limits, your monitoring control created a confidentiality gap.

Trusting an extension that still calls home. An assistant marketed as enterprise can still emit telemetry, license checks, or feature flags from the IDE side. In an air-gapped deployment that is not a nuance, it is a failed boundary. Verify at the packet level.

No change management for model updates. Bundles that arrive ad hoc, with no ticket, no signature check, and no evaluation gate, fail the change-management criterion even if the isolation is perfect. Make the crossing in step three the process.

Proving egress once. A go-live capture that is never repeated says nothing about the review period an auditor samples. Drift is the default; re-verification is the control.

Frequently asked questions

Does SOC 2 require an air-gapped deployment?

No. SOC 2 is about documented controls and the evidence that they operate, not network isolation. You can be SOC 2 compliant with plenty of internet-connected systems, as long as access, change management, and monitoring are in place and evidenced. Air-gapping is one strong control you can choose; it is not itself a SOC 2 requirement.

Does an air-gapped AI coding assistant automatically satisfy SOC 2?

No. Air-gapping makes several controls easy to evidence, especially confidentiality and logical access, because you can show there is no path out at all. But SOC 2 still requires documented access management, change management for model updates, monitoring, and the report from an independent auditor proving those controls operated over a period. Isolation removes a whole class of risk; it does not remove the paperwork.

Can AI coding assistants run fully air-gapped?

Yes, if every component runs offline: an open-weight model, a self-hosted inference server, an IDE extension that can point at your own endpoint with no mandatory outbound calls, and a retrieval index built inside the boundary. Model updates arrive as signed offline bundles brought across a one-way transfer rather than downloaded at runtime.

What is the difference between SOC 2 Type I and Type II for an AI deployment?

Type I attests that controls are designed appropriately at a single point in time. Type II attests that they operated effectively over a review period, typically three to twelve months, sampled by the auditor. Customers evaluating an AI coding assistant almost always ask for Type II, which is why standing evidence like access reviews and egress captures matters more than a one-time snapshot.

Which SOC 2 Trust Services Criteria apply to an AI coding assistant?

Security, the Common Criteria, is mandatory in every SOC 2 report. For a coding assistant, Confidentiality almost always applies because the prompts contain source code, and Availability usually applies because developers block on the tool. Processing Integrity and Privacy are added when scope calls for them, for example if the assistant handles regulated or personal data.

Do the assistant's prompt and completion logs fall under SOC 2 scope?

Yes. The moment you log prompts and completions for monitoring, those logs contain source code and sometimes secrets, so the log store itself becomes confidential data. It inherits the same access controls, retention limits, encryption, and monitoring as the rest of the system. An audit log is a SOC 2 asset for the Monitoring criterion and a SOC 2 liability for the Confidentiality criterion at the same time.

Where this fits at Numerata

Numerata is the infrastructure that makes this deployment installable inside your own boundary rather than assembled from parts. Lupine provides the GPUs, in your private cloud or fully air-gapped inside your environment. P95 fine-tunes the model on your own repositories so the weights are yours and stay in the boundary. NinetyFive serves it fast enough that a completion path on your own network beats a public endpoint on latency, with the gateway, redaction, and audit logging that Step 4 requires. It installs air-gapped, and the engagement includes the security review and the evidence set rather than leaving you to assemble them alone. If you are weighing this against a public copilot, we compare them honestly in should your coding assistant run on your own hardware, and the deployment mechanics are in how deployment works.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog