How to deploy AI coding assistants in air-gapped and SOC 2 environments
AUGUST 13, 2026 · NUMERATA TEAM
Deploying an AI coding assistant in an air-gapped or SOC 2
environment is two problems wearing one label. Air-gapping decides
where your source code is allowed to go. SOC 2 decides whether you
can prove your controls actually work. A deployment can satisfy one
and fail the other, so the first move is to stop treating them as the
same requirement. From there the process is six steps: separate the
two requirements, choose an assistant whose every component runs
offline, move the model across the air gap without opening a path
back, put a gateway in front for identity, redaction, and audit
logging, map each control to a SOC 2 Trust Services Criterion, and
prove zero egress with evidence you can hand an auditor. Here is each
step in detail.
Air-gapped and SOC 2 answer two different questions
These two words get bundled into one procurement line so often
that teams design for them as if they were a single control. They are
not. Air-gapping is a network fact: there is no path
in or out, verified rather than assumed. SOC 2 is an
attestation: an independent auditor examined your controls
and reported whether they were designed and operating effectively
against the Trust Services Criteria. One is a property of your
topology. The other is a property of your evidence.
The practical consequence is the thing most people get backwards.
You can hold a clean SOC 2 report while running services that talk to
the internet all day, because SOC 2 never asked you to disconnect,
only to control and prove. And you can air-gap a deployment perfectly
and still fail SOC 2, because pulling the plug says nothing about who
can access the box, how changes are approved, or whether anyone is
watching the logs. Air-gapping is a control you can choose. SOC 2 is
proof that your controls, whichever you chose, operate.
Design for both as one thing and you will over-build the network and under-build the evidence, or the reverse. They are separate requirements that happen to reinforce each other.
The good news is that they do reinforce each other. An air-gapped
coding assistant makes the hardest SOC 2 questions easy to answer,
because "does source code leave our boundary" stops being a policy
argument and becomes a packet capture. The rest of this guide builds
the deployment so that isolation and evidence come out of the same
architecture rather than being bolted on twice.
Step 1: Separate the air-gap requirement from the SOC 2 requirement
Before any hardware or model choice, write down which of the two
you are actually on the hook for, because the answer is often not
both. A regulator or a customer contract may demand air-gapping for a
specific class of code. A different customer may only demand a SOC 2
Type II report and never mention the network at all. Scope them
independently and you avoid the two expensive mistakes: air-gapping an
environment nobody required isolated, and assuming an isolated
environment is automatically audit-ready when it has no access reviews
behind it.
The distinction also decides who owns the work. Air-gapping is an
infrastructure decision, mostly settled once. SOC 2 is a continuous
program: controls that operate every day and generate evidence every
day, sampled by an auditor over a review period. If you are pursuing
both, treat the air gap as one control inside the SOC 2 program, not
as a substitute for it. We walk the isolation choice itself in more
depth in how to deploy
private AI; this post assumes you have landed on air-gapped and
focuses on doing it in a way an auditor can sign.
Step 2: Choose an assistant whose every component runs offline
An AI coding assistant is not one program. It is at least four, and
each has to run with no outbound calls for the deployment to be
genuinely air-gapped. The model weights and the inference server are
the parts everyone checks. The two that quietly break isolation are
the IDE extension and the retrieval index.
THE MODEL + SERVER
An open-weight model served by a self-hosted stack. The check: it loads weights from local disk, never fetches them at runtime, and its telemetry and update channels are off.
THE IDE EXTENSION
The usual failure point. Many extensions still send telemetry, license check-ins, or feature flags even in enterprise mode. You need one that points at your endpoint with no mandatory calls home.
THE RETRIEVAL INDEX
The context that makes the assistant good, built over your repos, docs, and tickets, lives inside the boundary. No hosted vector database, no embeddings API, no external reranker.
The IDE extension deserves the paranoia. A modern assistant does
not send "the file I am editing." It assembles context: neighboring
open tabs, imports resolved across the repository, symbol definitions
from the language server, sometimes terminal output and failing test
names. That whole bundle is what would leave if the extension had a
path out, which in an air-gapped deployment it must not. Instrument
the extension and read its requests before you trust the marketing;
the full inventory of where code escapes a network is in
every way source
code leaves, and how to close each path.
Step 3: Move the model across the air gap without opening a path back
An air-gapped model still has to get in somehow, and "somehow" is
exactly where isolation is usually lost. The wrong way is a temporary
firewall exception to pull weights or a package during install, which
is a path out that existed for an afternoon and will not appear in any
diagram. The right way is a one-way transfer: stage everything on a
connected machine, scan and sign it, carry the signed bundle across a
controlled transfer, and import it inside the boundary with signature
verification. Nothing about the import creates a return path.
The same one-way crossing that keeps the environment isolated is also your change-management evidence: a signature and an approval per model update, which is exactly what a SOC 2 auditor samples.
Notice what this step buys you twice. Operationally, it is how the
model gets in without ever letting anything out. For the audit, the
signature and the approval ticket on every crossing are the artifact
that answers "how do you control changes to production models," which
is a control an internet-connected deployment has to manufacture
separately. The air gap forced you to build change management as a
side effect. The full air-gapped workflow, including how updates keep
arriving on your schedule, is covered in
air-gapped LLM
training.
Step 4: Put a gateway in front for identity, redaction, and audit logging
A model server on its own is not auditable. It answers whoever can
reach it, redacts nothing, and remembers nothing. The component that
turns a private model into a reviewable system is a gateway that every
request passes through, and it is the single most load-bearing part of
a SOC 2 deployment. It is where identity, policy, redaction, and the
audit trail all live.
The model is the easy part. The gateway is what an auditor actually interviews: it is where you show identity, redaction, and a log, and it is the one place a control failure would show up.
Three of the gateway's jobs are worth stating plainly, because each
maps to a control an auditor will test. Identity and
access: requests authenticate against your existing directory
and are authorized per team, so "who could use the assistant, and on
which repositories" has an answer you can export. Redaction:
secrets and regulated fields are stripped before the prompt
reaches the model, because configuration files and fixtures routinely
carry credentials that no assistant should ingest. Audit
logging: an immutable record of requests, shipped to your
SIEM, which is both your monitoring evidence and, because it now holds
source code, a confidential store in its own right.
That last point is the subtle one. The moment you log prompts and
completions, the log contains code and possibly secrets, so the log
store inherits the same access controls, retention limits, and
encryption as production. An audit trail is a SOC 2 asset for
monitoring and a SOC 2 liability for confidentiality at the same time,
and the deployments that fail here are the ones that treated logging
as free. This is the same control surface a risk team probes in a
vendor review, which we lay out in
what your risk team will
ask.
Step 5: Map each control to a SOC 2 Trust Services Criterion
SOC 2 is built on five Trust Services Criteria: Security,
Availability, Processing Integrity, Confidentiality, and Privacy.
Security, known as the Common Criteria, is mandatory in every report.
For a coding assistant, Confidentiality almost always applies because
the prompts are source code, and Availability usually does because
developers block on the tool. Processing Integrity and Privacy come
into scope when the work calls for them. An audit does not grade your
architecture; it grades whether each control you claim is real and
operating. So the deliverable for this step is a table that ties every
control in the deployment to a criterion and, critically, to the
artifact that proves it.
Control in the deployment
SOC 2 criterion
Evidence you can produce
SSO and role-based access at the gateway
Logical access (CC6)
Directory group membership, periodic access-review exports
Secret and regulated-data redaction before the model
Confidentiality (C1)
Redaction rules in version control, sampled prompt logs
Signed, reviewed model bundles imported on a schedule
Change management (CC8)
Bundle signatures, change tickets, approval records
No outbound network path from the serving subnet
Confidentiality (C1)
Firewall deny-all rules, dated packet captures
Capacity and failover for the serving layer
Availability (A1)
Uptime monitoring, load-test results, runbooks
Evaluation gate before a new model serves traffic
Processing integrity (PI1)
Eval-suite results, sign-off before promotion
Read the table the way an auditor does, right column first. A
control with no producible evidence is not a control; it is an
intention. The reason an air-gapped deployment is comfortable to audit
is that its strongest evidence, the deny-all rule and the capture
showing nothing leaves, is objective and repeatable rather than a
matter of trust. You are not asking the auditor to believe code stays
inside. You are showing them the wall.
Step 6: Prove zero egress and keep the evidence auditors ask for
This is the step teams skip, and it is the one both the air-gap
claim and the SOC 2 report actually rest on. Absence of a complaint is
not absence of a connection. A deployment can satisfy every step above
and still have a path nobody intended: a dependency that phones home
on update, an analytics SDK bundled three layers deep, a container
that pulls a base image at restart. Prove the boundary directly, and
prove it more than once.
DENY BY DEFAULT
The serving subnet blocks outbound traffic as the default rule, not as an afterthought. Anything that needs to reach out has to be an explicit, reviewed exception, of which the target is zero.
CAPTURE AND KEEP
A packet capture over a real working session, showing no egress, dated and stored. This is the artifact that converts "we are air-gapped" from an assertion into evidence a Type II audit can sample.
RE-VERIFY ON CHANGE
Egress is not proven once at go-live. A model update, a dependency bump, or a config change can reopen a path, so the check runs again on every change and the result is filed.
The word that matters in SOC 2 Type II is period. Type I
says your controls were designed correctly on the day the auditor
looked. Type II says they operated correctly across a review window,
often three to twelve months, sampled without warning. That is why the
evidence in step five has to be standing rather than staged: access
reviews that happen on a cadence, captures that get re-taken on
change, tickets that exist because a process created them and not
because someone assembled a folder the week before the audit.
Isolation removes a class of risk in one architectural stroke; the
report is earned by keeping the evidence flowing.
Five mistakes that fail the audit
Treating air-gapping as if it were SOC 2, or the reverse.
The isolated environment with no access reviews and the compliant
environment with an unexamined outbound path are the same mistake from
opposite ends. Scope both, in step one.
Logging without putting the log in scope. Prompt
and completion logs contain source code and secrets. If the log store
does not inherit production access controls and retention limits, your
monitoring control created a confidentiality gap.
Trusting an extension that still calls home. An
assistant marketed as enterprise can still emit telemetry, license
checks, or feature flags from the IDE side. In an air-gapped
deployment that is not a nuance, it is a failed boundary. Verify at the
packet level.
No change management for model updates. Bundles
that arrive ad hoc, with no ticket, no signature check, and no
evaluation gate, fail the change-management criterion even if the
isolation is perfect. Make the crossing in step three the process.
Proving egress once. A go-live capture that is
never repeated says nothing about the review period an auditor
samples. Drift is the default; re-verification is the control.
Frequently asked questions
Does SOC 2 require an air-gapped deployment?
No. SOC 2 is about documented controls and the evidence that they operate, not network isolation. You can be SOC 2 compliant with plenty of internet-connected systems, as long as access, change management, and monitoring are in place and evidenced. Air-gapping is one strong control you can choose; it is not itself a SOC 2 requirement.
Does an air-gapped AI coding assistant automatically satisfy SOC 2?
No. Air-gapping makes several controls easy to evidence, especially confidentiality and logical access, because you can show there is no path out at all. But SOC 2 still requires documented access management, change management for model updates, monitoring, and the report from an independent auditor proving those controls operated over a period. Isolation removes a whole class of risk; it does not remove the paperwork.
Can AI coding assistants run fully air-gapped?
Yes, if every component runs offline: an open-weight model, a self-hosted inference server, an IDE extension that can point at your own endpoint with no mandatory outbound calls, and a retrieval index built inside the boundary. Model updates arrive as signed offline bundles brought across a one-way transfer rather than downloaded at runtime.
What is the difference between SOC 2 Type I and Type II for an AI deployment?
Type I attests that controls are designed appropriately at a single point in time. Type II attests that they operated effectively over a review period, typically three to twelve months, sampled by the auditor. Customers evaluating an AI coding assistant almost always ask for Type II, which is why standing evidence like access reviews and egress captures matters more than a one-time snapshot.
Which SOC 2 Trust Services Criteria apply to an AI coding assistant?
Security, the Common Criteria, is mandatory in every SOC 2 report. For a coding assistant, Confidentiality almost always applies because the prompts contain source code, and Availability usually applies because developers block on the tool. Processing Integrity and Privacy are added when scope calls for them, for example if the assistant handles regulated or personal data.
Do the assistant's prompt and completion logs fall under SOC 2 scope?
Yes. The moment you log prompts and completions for monitoring, those logs contain source code and sometimes secrets, so the log store itself becomes confidential data. It inherits the same access controls, retention limits, encryption, and monitoring as the rest of the system. An audit log is a SOC 2 asset for the Monitoring criterion and a SOC 2 liability for the Confidentiality criterion at the same time.
Where this fits at Numerata
Numerata is the infrastructure that makes this deployment
installable inside your own boundary rather than assembled from parts.
Lupine provides the GPUs, in your
private cloud or fully air-gapped inside your environment.
P95 fine-tunes the model on your own
repositories so the weights are yours and stay in the boundary.
NinetyFive serves it fast enough that a
completion path on your own network beats a public endpoint on
latency, with the gateway, redaction, and audit logging that Step 4
requires. It installs air-gapped, and the engagement includes the
security review and the evidence set rather than leaving you to
assemble them alone. If you are weighing this against a public copilot,
we compare them honestly in
should your coding
assistant run on your own hardware, and the deployment mechanics
are in how deployment works.