GPU orchestration is the software layer that decides which work
runs on which GPU, and when. Instead of pinning each job to a fixed
card, an orchestrator pools many GPUs across many machines, schedules
training and inference workloads onto them, shares or splits cards
between jobs, queues work that does not fit yet, and scales the pool
up or down as demand changes. Its job is to keep expensive hardware
busy and out of anyone's way.
If a single GPU is a machine, GPU orchestration is the operating
system for a room full of them. Without it, GPUs get handed out like
parking spaces: one team, one card, reserved whether or not anything
is running on it. With it, the whole fleet becomes one pool that a
scheduler fills on demand.
The problem: idle GPUs are the expensive default
A modern accelerator can cost as much as a car, and it costs that
whether it is training a model or sitting at zero percent. The natural
way teams acquire GPUs, one project reserving its own, produces a fleet
that is mostly idle: research runs in bursts, inference traffic is
peaky, and no single team keeps its own cards busy around the clock.
The result is a room full of expensive hardware at low average
utilization, while someone down the hall is queued up waiting for a
card. Orchestration exists to close that gap.
Pinned to desks, most of the fleet sits idle. Pooled behind a scheduler, the same GPUs stay busy.
What a GPU orchestrator actually does
Underneath the one-line definition, an orchestrator is doing four
jobs at once, continuously, as work arrives and finishes.
SCHEDULE & PLACE
Decide which GPU each job lands on, accounting for how much memory it needs and, for multi-GPU jobs, which cards are wired together fast enough to train across.
POOL & SHARE
Present scattered GPUs as one set of resources, and let more than one small job share a single card so a light inference model does not hold a whole GPU hostage.
QUEUE & PRIORITIZE
When demand exceeds the pool, hold jobs in a queue and order them by priority and fair-share, so a spike waits gracefully instead of failing or starving other teams.
SCALE & RECLAIM
Grow the pool when work backs up, and shrink it, in the extreme all the way to zero, when cards fall idle, so nothing sits powered and billing for traffic that is not there.
Orchestration vs. provisioning vs. a plain scheduler
Three words get used interchangeably and should not be.
Provisioning is making GPUs available in the first
place: acquiring the hardware or cloud instances and installing drivers.
A scheduler is the narrower engine that picks what runs
next on already-available resources. Orchestration is
the whole ongoing system around that engine: pooling GPUs across
machines, sharing cards, queuing, autoscaling, and recovering from
failures. Provisioning is a setup step you do occasionally; a scheduler
is one component; orchestration is the running discipline that keeps a
provisioned fleet actually utilized.
The scheduler spans a job across two cards, packs two half-jobs onto one, and holds the batch job in a queue until two cards free up.
Training and inference are two different scheduling problems
A good orchestrator has to serve two workloads that behave nothing
alike. Training is long-running and batch-like: a job asks for several
GPUs wired closely together, runs for hours or days, and cares about
throughput more than instant start. Inference is the opposite: many
short requests, bursty and latency-sensitive, where a queued request
is a slow response a user feels. Run them on separate fleets and both
sit idle half the time. Run them on one orchestrated pool and they fill
each other's gaps: inference gets priority when traffic spikes, and the
capacity it is not using flows back to training in between. Their peaks
rarely line up, which is exactly what makes sharing pay off.
Why it matters most when the GPUs are yours
In the public cloud, you can paper over weak orchestration with a
credit card: spin up more GPUs at the peak, hand them back after. A
private or air-gapped fleet cannot do that. You paid for those cards up
front, and they are all you have until the next purchase. That makes
orchestration the difference between the capacity you bought and the
capacity you can actually use. On a fixed fleet, every idle GPU is
money already spent and wasted, and every well-packed one is effective
capacity you did not have to buy again.
Manual GPU assignment vs. an orchestrated pool
The same fleet, run two ways.
Manual assignment
Orchestrated pool
Who picks the GPU
A person, by hand
The scheduler, automatically
Typical utilization
Low, cards sit idle
High, idle cards get reused
A sudden spike in demand
Wait, or hunt for a free card
Queued and placed as GPUs free up
Multiple teams sharing
First come, contention
Fair-share and priorities
Sharing one card
One job per GPU
Small jobs can share a GPU
Idle GPUs
Still powered, still costing
Scaled down, or to zero
Where this fits at Numerata
Lupine is Numerata's compute layer,
and this is the job it does. A coordinator aggregates scattered GPUs,
on-prem or in the cloud, into one pool exposed as a single device list,
reached through a drop-in CUDA shim so existing PyTorch, TensorFlow, and
JAX code runs against the pool unmodified. It scales to zero when idle,
so a fixed fleet is not sitting powered between research sprints. That
same pool feeds both sides of the stack:
P95 training jobs and
NinetyFive inference draw from one set of
GPUs, which is what keeps a private fleet near the utilization that
justifies buying it. If you are still working out how many cards you
need in the first place, our guide to
GPU sizing for LLM
inference is the companion to this one, and
what an inference
engine is covers the layer the pool serves.