Fine-tuning is the process of taking a model that's already been
pretrained on a broad, general dataset and continuing its training on
a smaller, task-specific dataset, so its weights shift toward your
domain, your tone, or your exact output format. The pretrained model
already knows language, reasoning, and general world knowledge.
Fine-tuning teaches it your particular problem.
Fine-tuning versus pretraining
Pretraining is where nearly all of a model's knowledge comes from:
a huge, general corpus, an enormous amount of compute, and weeks of
training to produce a base model that can write, reason, and follow
instructions reasonably well across almost any topic. Fine-tuning
starts from that finished base model instead of from scratch, and
continues training on a dataset that's orders of magnitude smaller,
built specifically for your task.
Pretraining builds general capability. Fine-tuning redirects it toward your task.
Fine-tuning versus prompting and RAG
Prompting and retrieval-augmented generation (RAG) are often the
first thing people reach for, and for good reason: neither requires
training a model at all. But they solve a different problem than
fine-tuning does.
PROMPTING
Instructions and examples go in the prompt itself, at inference time. Fast to iterate, no training required, but every request pays the token cost, and the model's weights never actually change.
RAG
Retrieves relevant documents and injects them into the prompt automatically. Great for keeping answers grounded in your data, but it's still steering a general model at inference time, not specializing it.
FINE-TUNING
Updates the model's weights directly, so the adaptation is baked in permanently. Worth it when you need consistent tone, a structured output format, or lower latency than a long prompt allows.
What actually happens during fine-tuning
You start from a pretrained checkpoint rather than random weights,
and continue training on a curated set of input and output examples
that demonstrate the behavior you want. Backpropagation adjusts the
model's weights toward those examples, but for a tiny fraction of the
steps pretraining took, since the model isn't learning language from
nothing, it's adjusting what it already knows.
Full fine-tuning updates every weight in the model, which is
accurate but expensive to store and serve, since each fine-tuned
variant is a full copy of the model. Parameter-efficient methods like
LoRA freeze the base model entirely and train a small set of added
weights instead, which is dramatically cheaper to train and to store,
at a modest cost to how much the model's behavior can shift.
Only the small adapter block is trainable in LoRA. Everything gray stays frozen.
When fine-tuning is worth it
Fine-tuning earns its cost when a prompt can't reliably get you
there. That usually shows up as needing consistent domain-specific
terminology or tone across every response, a structured output format
that has to hold up under edge cases rather than most of the time, or
latency and cost budgets that a long few-shot prompt would blow on
every single request. It's also the right tool when the behavior
you're teaching depends on proprietary data you don't want sitting in
a third-party prompt at all.
This is what P95,
Numerata's training layer, is built for: fine-tuning that runs on
infrastructure you already control, private cloud or fully
air-gapped, so the dataset that makes your model actually yours never
has to leave your environment to train it.