Standing up a private model is the easy part. The part that adds
three months is the security review, and it usually stalls not because
the answers are bad but because nobody assembled them before the
questionnaire arrived.
Below is the question set, grouped the way review boards actually
group it, with what each question is really asking and what a complete
answer contains. It is written to be forwarded: if you are the person
championing a deployment internally, send this to whoever runs third
party risk and ask which sections apply.
Why on-prem reviews go sideways
Most vendor risk frameworks were written for SaaS. They assume the
vendor holds the data, so they ask about the vendor's certifications,
the vendor's subprocessors, the vendor's breach history. An on-prem
deployment inverts every one of those assumptions, and reviewers
reasonably do not trust an inversion they have not seen before.
The fastest reviews are the ones where you demonstrate the
inversion rather than assert it. "No data leaves" is a claim. A packet
capture over a week showing zero outbound connections is evidence, and
evidence is what closes a review.
Each area closes on an artifact, not an assurance. Assemble the right column before the review starts.
1. Data flow and egress
The core of the review. Everything else is secondary if this
section does not hold up.
Question
What they are really asking
A complete answer
Where is data processed?
Does anything leave our network boundary
A diagram with the boundary drawn, showing every component inside it
What outbound connections exist?
Telemetry, licence checks, update pulls
An itemized list, ideally empty, plus a capture proving it
Is inference data logged, and where?
Whether prompts containing client data persist
Log destinations, retention, and how to turn content logging off
Does the vendor receive usage data?
Can you profile our activity
Yes or no. If yes, exactly which fields
How are model updates delivered?
Is there a network path in
Offline bundles, verified by signature, applied on your schedule
What happens on a network partition?
Does it fail closed or degrade
Behavior when isolated, ideally unchanged
2. Supply chain and provenance
Increasingly the section that takes longest, because AI stacks pull
large dependency trees and reviewers have learned to look.
Privilege model per component, and what can be dropped
3. Access and identity
Straightforward, and the section most likely to be answered
loosely. Reviewers notice.
Question
What they are really asking
A complete answer
Does the vendor have standing access?
Can someone outside reach production
None by default. If support access exists, how it is requested and logged
How does it integrate with our IdP?
Will this become a shadow user directory
SSO and SCIM support, group to role mapping
What is in the audit log?
Can we reconstruct who did what
Sample records, retention, export path to your SIEM
How are secrets handled?
Credentials in config files
Secret backend integration, and confirmation nothing is stored in plaintext
4. Model-specific risk
The newest section, and the one where generic questionnaires are
weakest. In regulated firms this is also where model risk management
obligations attach, so answer it as though a validation team will read
it, because one will.
Question
What they are really asking
A complete answer
What data was it trained on?
Lineage, and whether anything improper got in
Dataset inventory, source, and preprocessing for every fine-tune
Can we reproduce a given model version?
Auditability
Versioned data, code, config, and seed for each artifact
How do you know a new version is better?
Evaluation rigor before promotion
Held-out evaluation set, the metric, and the promotion threshold
How do you roll back?
Recovery when a version regresses
Retained prior versions and the time to revert
What is the prompt injection surface?
Can untrusted input change behavior
Which inputs are untrusted, and what the model is permitted to do
5. Continuity and exit
Asked last, and worth answering well, because a confident exit
answer removes the main objection to depending on anyone.
Question
What they are really asking
A complete answer
What if the vendor disappears?
Concentration risk
What keeps running unattended, and for how long
Who owns the trained model?
Is our tuned model an asset or a rental
Unambiguous ownership of weights and artifacts, in writing
Can we export everything?
Lock-in
Format of weights, data, and pipelines on export
Is source escrow available?
Continuity for a critical system
Yes or no, and the release conditions
What actually slows reviews down
Answering a claim with a claim. "The system does
not egress data" invites a follow-up. The same sentence with a capture
attached ends the thread.
Sending the SaaS questionnaire back unmodified.
Half the questions will not apply, and answering "N/A" twenty times
reads as evasion. Say why the framework does not fit, then answer the
spirit of each question.
Leaving model risk to the security team. Sections
one to three are security's remit. Section four usually belongs to
model validation or compliance, and if you do not route it there
deliberately it will surface late as a new objection.
Discovering the exit question at the end. Ownership
and export terms are contract questions with long internal lead times.
Raise them in week one.
Where this fits at Numerata
Security review is a named step in a
Numerata deployment, before install
rather than after it, and we bring the artifacts in the right-hand
column to it rather than waiting to be asked. On the answers above:
there are no outbound connections by construction, nothing phones
home, we hold no standing access to your environment, and the trained
weights are yours and stay in your environment whatever happens to the
commercial relationship. If your reviewers have a question this list
does not cover, we would genuinely like to see it.