A trading desk's model inputs are its edge: order flow, position
data, fill history, the exact logic behind a strategy. None of that
is data you can hand to a third-party API, and unlike most
back-office use cases, trading has a second constraint stacked on
top of privacy: latency. A network hop to a shared model endpoint
doesn't just risk exposing proprietary data, it can blow the timing
budget a strategy depends on entirely.
Here are six places private AI shows up on and around the desk.
Strategy logic goes in, a fine-tuned model comes out. The edge never leaves the building.
The six use cases
SIGNAL GENERATION
Fine-tune on your own order flow and market data to generate trading signals, so the model that finds your edge doesn't also hand a copy of it to a vendor.
EXECUTION & ORDER ROUTING
Decide how to slice and route orders in real time, on inference fast enough to run inside the same timing budget as the execution logic itself.
RISK & EXPOSURE MONITORING
Score portfolio risk continuously against your own book, using a model trained on your actual positions and history, not a generic risk proxy.
TRADE SURVEILLANCE
Detect spoofing, layering, and other market-abuse patterns in your own order and execution data, without shipping trading activity off-site to find them.
MARKET MICROSTRUCTURE ANALYSIS
Model order book dynamics and tick-level behavior for a specific venue or instrument, trained on the depth of history only your own capture provides.
RESEARCH & BACKTESTING COPILOTS
Summarize backtest runs and strategy research in-house, grounded in your own historical results instead of general market commentary.
Signal generation
A signal model is only as good as what it's trained on, and the
data that actually produces edge, your own order flow, fills, and
market microstructure history, is precisely the data you can't
route through a shared model API without handing a competitor's
lens straight into your strategy. Fine-tuning privately means the
model that finds your signal never has to leave the systems that
generated the data in the first place.
Execution and order routing
Deciding how to slice an order and where to route it has its own
timing budget, and that budget is usually measured in the same
units as the trade itself. Running that decision through a
third-party endpoint adds a network round trip on top of inference
time, which is often the difference between a routing model that's
useful and one that's already stale by the time it responds.
Risk and exposure monitoring
Generic risk models are built to be reasonable across almost any
portfolio, which means they're tuned for nobody's book in
particular. A model fine-tuned on your own positions and trading
history learns what exposure actually looks like for your specific
book, and can flag a concentration or a correlated risk that a
one-size-fits-all model would miss entirely.
Every hop outside your own infrastructure is time the strategy's timing budget has to absorb.
Trade surveillance
Spoofing, layering, and other market-abuse patterns show up in
the shape of your own order and execution data, which means
catching them requires a model trained on that exact data. Shipping
trading activity to a third party to run surveillance defeats the
purpose of keeping the activity confidential in the first place,
and most compliance regimes won't let you do it anyway.
Market microstructure analysis
Order book dynamics differ by venue, by instrument, and over
time, and modeling them well depends on the depth and specificity
of the history you train on. A model fine-tuned on your own
captured tick data learns the microstructure quirks of the exact
venues and instruments you actually trade, rather than an average
across markets you don't.
Research and backtesting copilots
Summarizing a backtest run, comparing it against prior strategy
iterations, or writing up research findings is exactly the kind of
task a language model is good at, provided it's grounded in your
own historical results rather than public market commentary. A
model fine-tuned on a firm's own research and backtest archive
picks up that context automatically. This is the one most desks
start with, so we've built it out end to end in a separate post:
an AI agent for
financial research, from retrieval and tool design through to
what the desk has to log.
Where this fits at Numerata
Every one of these starts the same way: fine-tune a model on
your own trading data with P95, on
Lupine compute that scales to zero
when you're not using it, and serve it through
NinetyFive, Numerata's inference
engine, at sub-50ms latency inside your own trust boundary, close
enough to the execution path to actually be useful. Private cloud
or fully air-gapped, the data that makes each of these models
useful never has to leave your infrastructure to train or run
it.