BLOG

Fine-tuning LLMs
for finance.

Fine-tuning a language model for finance means adapting a pretrained model's weights on domain-specific financial text and tasks, earnings calls, SEC filings, research notes, trade communications, so it produces more accurate, consistent, and compliant outputs than a general-purpose model prompted with instructions alone. For most teams, this means parameter-efficient fine-tuning (LoRA or QLoRA) on an open-weight base model, not full fine-tuning from scratch. The rest of this guide walks through when fine-tuning beats prompting or RAG, which open-source tools to use, and three deployable examples with code.

Why fine-tune instead of prompting or RAG?

Retrieval-augmented generation (RAG) and few-shot prompting solve knowledge gaps: giving a model the right context at inference time. Fine-tuning solves behavior gaps: teaching a model to reliably follow a specific output format, apply domain-specific judgment, or perform a narrow task with a much smaller and cheaper model. In finance specifically, fine-tuning tends to win in a few recurring situations.

NARROW, HIGH-VOLUME TASKS

Classifying transaction descriptions, tagging sentiment on earnings call transcripts, extracting structured fields from filings. A fine-tuned 7-8B model can match or beat a much larger general-purpose model at a fraction of the inference cost.

EXACT OUTPUT FORMATS

JSON schemas for downstream systems, specific taxonomy labels, and numeric formatting conventions all need to hold up on every single response, not most of them.

DATA THAT CAN'T LEAVE

Many trading firms and asset managers require on-prem or air-gapped inference, which rules out calling a hosted frontier model API for anything touching non-public information.

CONSISTENCY AT SCALE

Prompting can drift across a long batch job. A fine-tuned model's behavior is baked into the weights, so it doesn't wander over thousands of documents.

RAG still matters for grounding a model in current, firm-specific documents, a fund's own memos, a live position book. The two aren't mutually exclusive: a common production pattern is a fine-tuned model for task behavior, sitting on top of a RAG layer for facts.

Knowledge gaps versus behavior gaps in production Prompting and RAG fix knowledge gaps by supplying context at inference time. Fine-tuning fixes behavior gaps by changing the model's weights. Both feed into the same production system rather than competing. PROMPTING & RAG FIXES KNOWLEDGE GAPS FINE-TUNING FIXES BEHAVIOR GAPS PRODUCTION SYSTEM
Not competing approaches. Most production systems need both, for different reasons.

Core techniques, briefly

TechniqueWhat it doesWhen to use it
Full fine-tuningUpdates all model weightsRare outside large labs, expensive, needs large curated datasets
LoRATrains small low-rank adapter matrices, freezes base weightsDefault choice for most teams, cheap, fast, easy to swap adapters per task
QLoRALoRA on a 4-bit quantized base modelSame as LoRA but fits on a single consumer or single-datacenter GPU
Instruction tuningFine-tunes on (instruction, response) pairsTeaching a model a task format, e.g. extract these five fields as JSON
DPO / RLHFAligns outputs to preference dataRefining tone or judgment calls, less common for structured financial tasks

For nearly all financial NLP use cases, classification, extraction, structured summarization, QLoRA on an open-weight 7-8B model is the practical starting point. It runs on a single GPU, trains in hours, and gets you most of the accuracy gain full fine-tuning would deliver.

Open-source tools worth knowing

ToolBest forNotes
Hugging Face transformers + PEFT + TRLFull control, custom training loopsThe foundational libraries everything else builds on
UnslothFastest QLoRA training on a single GPU2x+ speedups and lower VRAM vs. vanilla PEFT, good for iterating quickly
AxolotlConfig-driven fine-tuning at scaleYAML config instead of code, easy to version-control training runs
LLaMA-FactoryGUI + CLI, broad model supportGood if you want a UI for less technical team members
vLLMServing fine-tuned models in productionHigh-throughput inference server, supports LoRA adapter hot-swapping
llama.cpp / GGUFOn-prem, CPU, or air-gapped inferenceQuantize and run without a GPU cluster, relevant for compliance-constrained deployments

Example 1: Financial sentiment classification with QLoRA + Unsloth

A common first project: classify sentiment (positive, negative, neutral) on financial statements, earnings guidance, analyst notes, headlines, more accurately than a generic sentiment model, which tends to misread finance-specific phrasing (e.g. "beat lowered guidance" reading as negative when it's actually a positive surprise).

Dataset: Financial PhraseBank (public, ~5,000 labeled sentences from financial news) or FiQA sentiment data, both open and Hugging Face-hosted.

Base model: Llama 3.1 8B or Mistral 7B, both permissively licensed for commercial use.

PYTHON

from unsloth import FastLanguageModel
from datasets import load_dataset
from trl import SFTTrainer
from transformers import TrainingArguments

# Load base model in 4-bit for QLoRA
model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/Meta-Llama-3.1-8B",
    max_seq_length=2048,
    load_in_4bit=True,
)

model = FastLanguageModel.get_peft_model(
    model,
    r=16,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    lora_alpha=16,
    lora_dropout=0,
    bias="none",
)

dataset = load_dataset("financial_phrasebank", "sentences_allagree")

def format_example(row):
    return {
        "text": f"### Statement:\n{row['sentence']}\n\n### Sentiment:\n{row['label']}"
    }

dataset = dataset["train"].map(format_example)

trainer = SFTTrainer(
    model=model,
    tokenizer=tokenizer,
    train_dataset=dataset,
    dataset_text_field="text",
    max_seq_length=2048,
    args=TrainingArguments(
        per_device_train_batch_size=4,
        gradient_accumulation_steps=4,
        num_train_epochs=3,
        learning_rate=2e-4,
        fp16=True,
        output_dir="outputs",
    ),
)
trainer.train()

This runs comfortably on a single 24GB consumer GPU. Once trained, save just the LoRA adapter (a few hundred MB) rather than a full model checkpoint, this is what makes iterating on multiple financial tasks cheap: one base model, many small task-specific adapters.

One base model, several hot-swappable task adapters A single frozen base model with three small LoRA adapters attached for different financial tasks: sentiment classification, entity extraction, and financial QA. Each adapter is a few hundred megabytes and can be swapped in independently. BASE MODEL (FROZEN) SENTIMENT ADAPTER EXTRACTION ADAPTER FINANCIAL QA ADAPTER
Each adapter is a few hundred megabytes. Swap tasks without reloading the base model.

Example 2: Structured entity extraction from filings with Axolotl

Extracting structured fields from SEC filings or credit agreements (counterparties, dollar amounts, covenant thresholds, dates) is a higher-value, harder task than sentiment. It benefits from instruction tuning on document excerpt to JSON pairs.

Axolotl config for this looks like:

AXOLOTL.YML

base_model: mistralai/Mistral-7B-v0.3
load_in_4bit: true
adapter: qlora

datasets:
  - path: ./data/filings_extraction.jsonl
    type: alpaca

lora_r: 32
lora_alpha: 16
lora_dropout: 0.05
lora_target_modules:
  - q_proj
  - v_proj

sequence_len: 4096
gradient_accumulation_steps: 4
micro_batch_size: 2
num_epochs: 3
learning_rate: 0.0002

Your filings_extraction.jsonl would contain examples like:

JSON

{
  "instruction": "Extract the borrower, lender, and total facility amount from this excerpt as JSON.",
  "input": "This Credit Agreement is entered into by and between Acme Capital LLC ('Borrower') and Northstar Lending Partners ('Lender') for a total facility of $75,000,000...",
  "output": "{\"borrower\": \"Acme Capital LLC\", \"lender\": \"Northstar Lending Partners\", \"facility_amount\": 75000000}"
}

Longer sequence length (4096 or more) matters here since filing excerpts run long, a case where an efficient attention implementation, FlashAttention, which both Axolotl and Unsloth support out of the box, meaningfully affects whether training is practical on a single GPU.

Example 3: Financial QA with LLaMA-Factory

For question answering over financial documents, "What was the year-over-year change in operating margin?", the FinQA and ConvFinQA datasets are the standard open benchmarks, combining tabular financial data with multi-step numerical reasoning.

LLaMA-Factory exposes this as a CLI without writing training code:

BASH

llamafactory-cli train \
  --stage sft \
  --model_name_or_path Qwen/Qwen2.5-7B-Instruct \
  --dataset finqa \
  --finetuning_type lora \
  --lora_target q_proj,v_proj \
  --output_dir ./finqa-lora \
  --per_device_train_batch_size 4 \
  --learning_rate 2e-4 \
  --num_train_epochs 3 \
  --quantization_bit 4

Numerical reasoning tasks are worth evaluating carefully post-training: measure exact-match accuracy on held-out FinQA questions, not just training loss, since fine-tuned models can look fluent while still getting arithmetic wrong. It's common to see fine-tuning lift task-specific accuracy dramatically. Base models scoring in the 30-40% range on financial text classification tasks are a common enough starting point, with fine-tuned versions of the same model reaching 80%+ on the same eval, though the size of the gain depends heavily on how narrow and well-labeled the target task is.

Deployment: getting fine-tuned models into production

VLLM

High-throughput serving with native support for hot-swapping LoRA adapters, one base model process serving multiple fine-tuned task adapters.

LLAMA.CPP / GGUF

Quantized inference for on-prem or air-gapped deployment with no external network dependency, a common requirement for trading and asset management firms handling non-public information.

EVALUATE BEFORE ROLLOUT

Hold out a labeled test set and measure task-specific metrics, F1 for classification, exact-match for extraction and QA, not generic benchmarks that don't reflect your actual task.

FAQ

What's the difference between fine-tuning and RAG for financial LLMs?

RAG retrieves relevant documents at inference time to ground a model's answers in current, firm-specific facts. Fine-tuning changes the model's underlying behavior: output format, task accuracy, domain judgment. Production systems often use both together.

How much data do I need to fine-tune an LLM for a financial task?

For LoRA or QLoRA on a narrow classification or extraction task, a few hundred to a few thousand well-labeled examples is often enough to see a meaningful accuracy improvement over the base model. Broader instruction-following or reasoning tasks need more.

Can financial LLMs be fine-tuned entirely on-premise?

Yes. QLoRA training and llama.cpp/GGUF inference can run on a single GPU or on-prem cluster with no external API calls, which is why parameter-efficient fine-tuning of open-weight models is the standard approach for firms with data residency or air-gap requirements.

Which open-source base models are commonly used for financial fine-tuning?

Llama 3.1, Mistral, and Qwen2.5 are the most common permissively licensed base models. FinGPT also publishes financial-domain-adapted open models that can serve as a starting point instead of a general-purpose base model.

Does fine-tuning reduce hallucination in financial outputs?

It can reduce format and task-specific errors such as wrong labels or malformed JSON, but it does not eliminate factual hallucination on numerical reasoning. Numerical outputs should still be validated against source documents, especially for anything used in compliance or client-facing contexts.

Where this fits at Numerata

Everything above runs the same way on P95, Numerata's training layer: fine-tune on your own filings, transcripts, or trade data with the same code running locally, in a sweep, or as a cloud job, on infrastructure you control, private cloud or fully air-gapped. NinetyFive then serves the resulting adapters at sub-50ms latency, so the LoRA hot-swapping pattern above isn't just a demo, it's how production traffic actually gets routed. For the broader case on why finance is one of the industries this matters most for, see AI for finance.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog