Fine-tuning a language model for finance means adapting a
pretrained model's weights on domain-specific financial text and
tasks, earnings calls, SEC filings, research notes, trade
communications, so it produces more accurate, consistent, and
compliant outputs than a general-purpose model prompted with
instructions alone. For most teams, this means parameter-efficient
fine-tuning (LoRA or QLoRA) on an open-weight base model, not full
fine-tuning from scratch. The rest of this guide walks through when
fine-tuning beats prompting or RAG, which open-source tools to use,
and three deployable examples with code.
Why fine-tune instead of prompting or RAG?
Retrieval-augmented generation (RAG) and few-shot prompting solve
knowledge gaps: giving a model the right context at inference time.
Fine-tuning solves behavior gaps: teaching a model to reliably
follow a specific output format, apply domain-specific judgment, or
perform a narrow task with a much smaller and cheaper model. In
finance specifically, fine-tuning tends to win in a few recurring
situations.
NARROW, HIGH-VOLUME TASKS
Classifying transaction descriptions, tagging sentiment on earnings call transcripts, extracting structured fields from filings. A fine-tuned 7-8B model can match or beat a much larger general-purpose model at a fraction of the inference cost.
EXACT OUTPUT FORMATS
JSON schemas for downstream systems, specific taxonomy labels, and numeric formatting conventions all need to hold up on every single response, not most of them.
DATA THAT CAN'T LEAVE
Many trading firms and asset managers require on-prem or air-gapped inference, which rules out calling a hosted frontier model API for anything touching non-public information.
CONSISTENCY AT SCALE
Prompting can drift across a long batch job. A fine-tuned model's behavior is baked into the weights, so it doesn't wander over thousands of documents.
RAG still matters for grounding a model in current,
firm-specific documents, a fund's own memos, a live position book.
The two aren't mutually exclusive: a common production pattern is a
fine-tuned model for task behavior, sitting on top of a RAG layer
for facts.
Not competing approaches. Most production systems need both, for different reasons.
Core techniques, briefly
Technique
What it does
When to use it
Full fine-tuning
Updates all model weights
Rare outside large labs, expensive, needs large curated datasets
LoRA
Trains small low-rank adapter matrices, freezes base weights
Default choice for most teams, cheap, fast, easy to swap adapters per task
QLoRA
LoRA on a 4-bit quantized base model
Same as LoRA but fits on a single consumer or single-datacenter GPU
Instruction tuning
Fine-tunes on (instruction, response) pairs
Teaching a model a task format, e.g. extract these five fields as JSON
DPO / RLHF
Aligns outputs to preference data
Refining tone or judgment calls, less common for structured financial tasks
For nearly all financial NLP use cases, classification,
extraction, structured summarization, QLoRA on an
open-weight 7-8B model is the practical starting point. It
runs on a single GPU, trains in hours, and gets you most of the
accuracy gain full fine-tuning would deliver.
Open-source tools worth knowing
Tool
Best for
Notes
Hugging Face transformers + PEFT + TRL
Full control, custom training loops
The foundational libraries everything else builds on
Unsloth
Fastest QLoRA training on a single GPU
2x+ speedups and lower VRAM vs. vanilla PEFT, good for iterating quickly
Axolotl
Config-driven fine-tuning at scale
YAML config instead of code, easy to version-control training runs
LLaMA-Factory
GUI + CLI, broad model support
Good if you want a UI for less technical team members
Quantize and run without a GPU cluster, relevant for compliance-constrained deployments
Example 1: Financial sentiment classification with QLoRA + Unsloth
A common first project: classify sentiment (positive, negative,
neutral) on financial statements, earnings guidance, analyst notes,
headlines, more accurately than a generic sentiment model, which
tends to misread finance-specific phrasing (e.g. "beat lowered
guidance" reading as negative when it's actually a positive
surprise).
Base model: Llama 3.1 8B or Mistral 7B, both
permissively licensed for commercial use.
PYTHON
from unsloth import FastLanguageModel
from datasets import load_dataset
from trl import SFTTrainer
from transformers import TrainingArguments
# Load base model in 4-bit for QLoRA
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Meta-Llama-3.1-8B",
max_seq_length=2048,
load_in_4bit=True,
)
model = FastLanguageModel.get_peft_model(
model,
r=16,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
lora_alpha=16,
lora_dropout=0,
bias="none",
)
dataset = load_dataset("financial_phrasebank", "sentences_allagree")
def format_example(row):
return {
"text": f"### Statement:\n{row['sentence']}\n\n### Sentiment:\n{row['label']}"
}
dataset = dataset["train"].map(format_example)
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
train_dataset=dataset,
dataset_text_field="text",
max_seq_length=2048,
args=TrainingArguments(
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
num_train_epochs=3,
learning_rate=2e-4,
fp16=True,
output_dir="outputs",
),
)
trainer.train()
This runs comfortably on a single 24GB consumer GPU. Once
trained, save just the LoRA adapter (a few hundred MB) rather than
a full model checkpoint, this is what makes iterating on multiple
financial tasks cheap: one base model, many small task-specific
adapters.
Each adapter is a few hundred megabytes. Swap tasks without reloading the base model.
Example 2: Structured entity extraction from filings with Axolotl
Extracting structured fields from SEC filings or credit
agreements (counterparties, dollar amounts, covenant thresholds,
dates) is a higher-value, harder task than sentiment. It benefits
from instruction tuning on document excerpt to JSON pairs.
Your filings_extraction.jsonl would contain examples
like:
JSON
{
"instruction": "Extract the borrower, lender, and total facility amount from this excerpt as JSON.",
"input": "This Credit Agreement is entered into by and between Acme Capital LLC ('Borrower') and Northstar Lending Partners ('Lender') for a total facility of $75,000,000...",
"output": "{\"borrower\": \"Acme Capital LLC\", \"lender\": \"Northstar Lending Partners\", \"facility_amount\": 75000000}"
}
Longer sequence length (4096 or more) matters here since filing
excerpts run long, a case where an efficient attention
implementation, FlashAttention, which both Axolotl and Unsloth
support out of the box, meaningfully affects whether training is
practical on a single GPU.
Example 3: Financial QA with LLaMA-Factory
For question answering over financial documents, "What was the
year-over-year change in operating margin?", the
FinQA
and ConvFinQA
datasets are the standard open benchmarks, combining tabular
financial data with multi-step numerical reasoning.
LLaMA-Factory exposes this as a CLI without writing training
code:
Numerical reasoning tasks are worth evaluating carefully
post-training: measure exact-match accuracy on held-out FinQA
questions, not just training loss, since fine-tuned models can look
fluent while still getting arithmetic wrong. It's common to see
fine-tuning lift task-specific accuracy dramatically. Base models
scoring in the 30-40% range on financial text classification tasks
are a common enough starting point, with fine-tuned versions of the
same model reaching 80%+ on the same eval, though the size of the
gain depends heavily on how narrow and well-labeled the target task
is.
Deployment: getting fine-tuned models into production
VLLM
High-throughput serving with native support for hot-swapping LoRA adapters, one base model process serving multiple fine-tuned task adapters.
LLAMA.CPP / GGUF
Quantized inference for on-prem or air-gapped deployment with no external network dependency, a common requirement for trading and asset management firms handling non-public information.
EVALUATE BEFORE ROLLOUT
Hold out a labeled test set and measure task-specific metrics, F1 for classification, exact-match for extraction and QA, not generic benchmarks that don't reflect your actual task.
FAQ
What's the difference between fine-tuning and RAG for financial LLMs?
RAG retrieves relevant documents at inference time to ground a model's answers in current, firm-specific facts. Fine-tuning changes the model's underlying behavior: output format, task accuracy, domain judgment. Production systems often use both together.
How much data do I need to fine-tune an LLM for a financial task?
For LoRA or QLoRA on a narrow classification or extraction task, a few hundred to a few thousand well-labeled examples is often enough to see a meaningful accuracy improvement over the base model. Broader instruction-following or reasoning tasks need more.
Can financial LLMs be fine-tuned entirely on-premise?
Yes. QLoRA training and llama.cpp/GGUF inference can run on a single GPU or on-prem cluster with no external API calls, which is why parameter-efficient fine-tuning of open-weight models is the standard approach for firms with data residency or air-gap requirements.
Which open-source base models are commonly used for financial fine-tuning?
Llama 3.1, Mistral, and Qwen2.5 are the most common permissively licensed base models. FinGPT also publishes financial-domain-adapted open models that can serve as a starting point instead of a general-purpose base model.
Does fine-tuning reduce hallucination in financial outputs?
It can reduce format and task-specific errors such as wrong labels or malformed JSON, but it does not eliminate factual hallucination on numerical reasoning. Numerical outputs should still be validated against source documents, especially for anything used in compliance or client-facing contexts.
Where this fits at Numerata
Everything above runs the same way on P95,
Numerata's training layer: fine-tune on your own filings, transcripts,
or trade data with the same code running locally, in a sweep, or as
a cloud job, on infrastructure you control, private cloud or fully
air-gapped. NinetyFive then serves the
resulting adapters at sub-50ms latency, so the LoRA hot-swapping
pattern above isn't just a demo, it's how production traffic
actually gets routed. For the broader case on why finance is one of
the industries this matters most for, see
AI for finance.