Financial text classifiers: fine-tune, don't prompt.
OCTOBER 8, 2026 · NUMERATA TEAM
A financial text classifier assigns each piece of financial text,
a news sentence, a client chat message, a compliance-flagged email,
to one label from a fixed set. The fastest way to build one today is
to ask a general-purpose LLM for the label. The most accurate way, by
a wide margin, is to fine-tune a small model on a few hundred labeled
examples. In Numerata's benchmark across seven public classification
tasks, a prompted 0.6B-parameter base model scored 29.9
F1 on Financial PhraseBank sentiment and 11.6
on 77-way Banking77 intent. A LoRA fine-tune of the identical model
scored 82.1 and 91.2, beat both
TF-IDF and full fine-tuning on six of seven tasks, and answered in
about 30 ms.
This post walks through the four standard ways to build a text
classifier, the full benchmark results, why each approach wins or
loses where it does, and a reproducible recipe for building your own.
+52.2 PTS VS. PROMPTING
Financial PhraseBank sentiment, same 0.6B model, same hardware: 29.9 F1 prompted, 82.1 after LoRA fine-tuning.
6 OF 7 TASKS WON
Against TF-IDF and full fine-tuning. The seventh, PhraseBank, is within 0.8 points of full tuning.
91% FEWER LABELS
To reach 95.5 macro-F1 on TREC-QC: 480 labeled examples, against 5,452 for full fine-tuning.
What is a financial text classifier?
Classification is the workhorse of financial NLP. It's less
glamorous than research agents or document Q&A, but it's where
most of the volume is: every headline that needs a sentiment score,
every client message that needs routing, every outbound
communication that needs a surveillance flag. The output is one label
from a closed set, so the task is easy to specify, easy to evaluate,
and very easy to run millions of times a day. That last property is
why model size and latency matter more here than almost anywhere
else.
The seven benchmark tasks below are public datasets, but each one
is a direct stand-in for a production workload on a trading desk or
in a bank's operations stack.
Benchmark
Classes
What it tests
Financial workload it stands in for
Financial PhraseBank
3
Sentiment of sentences from financial news
News and filings sentiment signals, earnings-call tone
Banking77
77
Fine-grained intent of retail banking queries
Client service routing, chatbot triage, ticket tagging
Query routing for research assistants and internal search
Four ways to build a text classifier
Every approach in the benchmark turns text into a label. They
differ in where the knowledge of your label set lives: in a prompt,
in word counts, or in the model's weights.
Prompted base model (zero-shot). Describe the
labels in a prompt and ask the model to answer with one. No
training at all. The label set lives entirely in the context window,
so the model has to read it on every call and then generate a
label as text, which your code has to parse.
TF-IDF + linear classifier. The classic
baseline. Represent each document as weighted word and n-gram
counts, then fit logistic regression or a linear SVM. Trains in
seconds on a CPU. It has no notion of word order or meaning beyond
the n-grams it sees, which is a strength on keyword-driven tasks
and a weakness on phrasing-driven ones.
Full fine-tuning. Update every weight in the
pretrained model on your labeled data. Maximum flexibility, but
with a few hundred examples and hundreds of millions of free
parameters, it overfits easily and can overwrite useful pretrained
knowledge.
LoRA fine-tuning, merged. Freeze the
pretrained weights and train a small low-rank update alongside
them. After training, fold the update back into the base weights so
the deployed model is a single ordinary checkpoint. This is the
approach labeled "Numerata" in the results.
Prompting keeps the label set in the context window. The other three put it somewhere the model can learn it.
Benchmark setup
All four approaches were run against the same seven tasks. The
three model-based approaches share one 0.6B-parameter base model, so
every difference in the results comes from how the model was adapted,
not from model size. Scores are F1 on each task's held-out test set,
higher is better (macro-F1 for TREC-QC). Latency is P50 at batch size
one. Every number was produced on a single NVIDIA GH200 inside a
self-hosted environment, with no data or model leaving it.
Results: accuracy across seven tasks
Task
Classes
Prompted base
TF-IDF
Full tune
Numerata (LoRA)
SMS Spam
2
0.0
90.2
94.6
97.1
AG News
4
67.0
88.4
91.6
92.0
Financial PhraseBank
3
29.9
66.5
82.9
82.1
Banking77
77
11.6
86.6
85.4
91.2
SpamAssassin
2
3.9
94.0
87.1
96.9
Emotion
6
29.5
81.4
86.9
89.1
TREC-QC
6
7.2
76.0
96.7
96.9
Mean of 7
21.3
83.3
89.3
92.2
Same tasks as the table. Each row spans worst to best approach; hover a marker for its exact score.
Three patterns stand out. Prompting is not a weaker version of
the other approaches, it's a different regime entirely, averaging
21.3 F1 against 83 to 92 for everything else. TF-IDF is a far
stronger baseline than its reputation suggests, and beats full
fine-tuning on two tasks. And the LoRA fine-tune is the only
approach that is never meaningfully behind on any task.
Why prompting a small model fails at classification
The prompted scores are not a rounding error. On Financial
PhraseBank, a three-way task, 29.9 F1 is roughly what you would get
by guessing. On SMS Spam the prompted model scores 0.0, and on
SpamAssassin 3.9: it essentially never catches the spam it was
asked to find. Three mechanisms explain most of this.
The label set has to fit in working memory.
For Banking77, the prompt has to enumerate 77 intents, many of them
near-synonyms ("card not working" vs. "card swallowed" vs. "declined
card payment"). A 0.6B model reading that list once per request has
no stable internal representation of the boundaries between them.
It scored 11.6. TF-IDF, which learned those boundaries from
examples, scored 86.6.
Base-model priors don't match your taxonomy.
Financial sentiment is not everyday sentiment. "Operating profit
fell less than expected" is positive news. "The company will
maintain its dividend" is usually neutral. A model that has never
been shown your labeling convention defaults to the everyday
reading.
Generated labels have to be parsed. A prompted
model returns free text. Any response that isn't exactly a valid
label, a hedge, an explanation, a label that isn't in the set,
counts as a miss. Small models are especially prone to this, and on
a binary task where the model drifts toward one answer, F1 on the
positive class collapses toward zero.
Larger hosted models prompt considerably better than a 0.6B model
does. But they cost far more per call, they're slower, and for most
financial text they mean sending client communications or
non-public information to a third party. Fine-tuning gets a small
model to the accuracy you need, inside your own environment.
When TF-IDF is good enough, and when it isn't
TF-IDF scored 94.0 on SpamAssassin, beating full fine-tuning by
6.9 points, and 86.6 on Banking77, beating it by 1.2. Both are
tasks where specific words carry most of the signal: spam has a
vocabulary, and banking intents are usually named by their nouns
("top-up", "PIN", "exchange rate").
It's weakest where meaning depends on how words combine. On
Financial PhraseBank it scored 66.5, more than 15 points behind both
fine-tuned models, because "profit rose" and "loss rose" share most of
their n-grams and mean opposite things. TREC-QC is similar: question
type depends on structure ("how many" vs. "how did"), and TF-IDF
trailed by about 21 points.
The practical rule: always train TF-IDF first. It
takes minutes, it tells you how much a model has to beat to earn its
serving cost, and on a keyword-driven task it may be all you
need.
Why LoRA beat full fine-tuning
LoRA (low-rank adaptation) freezes each pretrained weight matrix
W and learns an update ΔW = BA, where
B and A are thin matrices of rank
r, typically 8 to 64. For a 1,024 × 1,024
projection at rank 16, that's about 33,000 trainable parameters
instead of about a million. Across a whole model, the trainable
parameter count usually drops to around 1% of the total or less.
Train a small update, fold it into the weights, serve one ordinary checkpoint.
That constraint is the point. With a few hundred or a few
thousand labeled examples, full fine-tuning has far more freedom than
the data can pin down, so it tends to memorize the training set and
drift away from the general language knowledge that made the
pretrained model useful. The low-rank update acts as a strong
regularizer: the model can only move in a small number of directions,
so it keeps its pretrained representations and learns the decision
boundary on top of them.
The benchmark shows exactly where that matters. The gap between
LoRA and full tuning is widest on SpamAssassin (+9.8) and Banking77
(+5.8), where training data is limited relative to the variety in
the inputs or the size of the label space. It's narrowest on TREC-QC
(+0.2) and AG News (+0.4), where full tuning had enough data to
converge. On Financial PhraseBank full tuning edged ahead by 0.8, a
tie for practical purposes.
Data efficiency: how many labels do you need?
For most financial teams, labeled data is the real bottleneck.
Labels have to come from analysts or compliance staff whose time is
expensive, and sometimes the text itself can't be sent to an outside
labeling vendor. So the useful question is not "which approach is
best with unlimited data" but "which approach is best with the 200
examples we can realistically label this month".
To answer it, the benchmark trained each approach on TREC-QC at
five dataset sizes, from 60 labeled examples up to the full 5,452.
TREC-QC macro-F1. Labeled values are exact; unlabeled points are read off the benchmark chart and approximate.
At 120 labeled examples, the LoRA fine-tune
reached 83.6 macro-F1. Full fine-tuning on the same
120 examples scored about 48, and TF-IDF about 42: leads of 35.5 and
41.4 points. The LoRA run crossed 95.5 at 480
examples. Full fine-tuning needed the entire 5,452-example
training set to reach a comparable score (96.7). The LoRA run got
there with 91% fewer labels; put the other way, full
fine-tuning needed more than 11 times the labeling for the same
result.
That's the difference between a labeling project and a labeling
afternoon. 120 examples is something a single analyst can produce in
a few hours, and a dataset small enough to review line by line before
it ever touches a model, which matters when compliance has to sign
off on training data.
Latency: why the fine-tuned model is faster too
You might expect the more accurate model to be the slower one.
Here it's the opposite.
Same 0.6B base model, same GH200. The difference is entirely in what each request has to do.
The merged fine-tune runs at about 30 ms P50 at
batch size one. Prompting the same base model takes 63 to 174
ms depending on the task, 2 to 6 times slower. Two things
account for it:
Prompt length. A prompted classifier has to
process the instructions and the full label list on every request
before it reads the input. The more labels, the longer the prefill,
which is why the prompted range is so wide across tasks. A
fine-tuned classifier processes only the input text.
Output length. A prompted model generates a
label as one or more tokens, often with preamble. A fine-tuned
classifier produces a single prediction, one forward pass.
Merging matters for the same reason. Because the LoRA update is
folded into the base weights before deployment, the served model has
exactly the architecture and cost of the base model. There's no
adapter lookup or extra matrix multiply per layer. At 30 ms a
classifier can sit inline in a request path, scoring every headline
as it arrives or every message before it's routed, instead of running
as an overnight batch.
How to build a financial text classifier
The recipe below follows the benchmark's lessons. The code is a
minimal, standard Hugging Face transformers +
peft setup, one reasonable way to do it rather than the
exact benchmark harness.
Pin down the taxonomy. Write a one-line
definition and two examples for every label, including the
ambiguous cases. Most classifier failures in production are
taxonomy disagreements, not model errors.
Label a few hundred examples, stratified.
Start with 100 to 500, making sure rare classes are represented.
Hold out a separate test set of at least a few hundred that nobody
trains on or tunes against.
Train a TF-IDF baseline. It sets the bar the
model has to clear.
LoRA fine-tune a small model with a
classification head.
Evaluate with macro-F1 and a confusion
matrix, not accuracy alone.
Merge, serve, and monitor. Track the predicted
label distribution over time. A shift is usually the first sign of
input drift.
PYTHON · LORA CLASSIFIER
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import LoraConfig, TaskType, get_peft_model
BASE = "path/to/small-base-model" # any ~0.5-1B open-weight model
LABELS = ["negative", "neutral", "positive"]
tok = AutoTokenizer.from_pretrained(BASE)
if tok.pad_token is None:
tok.pad_token = tok.eos_token
model = AutoModelForSequenceClassification.from_pretrained(
BASE,
num_labels=len(LABELS),
id2label=dict(enumerate(LABELS)),
label2id={l: i for i, l in enumerate(LABELS)},
)
model.config.pad_token_id = tok.pad_token_id
# SEQ_CLS keeps the new classification head fully trainable;# only the low-rank updates are trained inside the transformer.
lora = LoraConfig(
task_type=TaskType.SEQ_CLS,
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)
model = get_peft_model(model, lora)
model.print_trainable_parameters()
# ... train with transformers.Trainer on your labeled split ...# Fold BA into W: one ordinary checkpoint, no adapter at serve time.
merged = model.merge_and_unload()
merged.save_pretrained("fin-sentiment-merged")
tok.save_pretrained("fin-sentiment-merged")
PYTHON · EVALUATION
from sklearn.metrics import classification_report, confusion_matrix, f1_score
# y_true, y_pred: label ids on the held-out test set.# For a prompted baseline, map unparseable outputs to a "wrong" id# rather than dropping them, or the comparison is flattering.
print("macro-F1:", f1_score(y_true, y_pred, average="macro"))
print(classification_report(y_true, y_pred, target_names=LABELS, digits=3))
print(confusion_matrix(y_true, y_pred))
A note on metrics. Financial label sets are almost always
imbalanced: most messages aren't spam, most news sentences are
neutral. A model that always predicts the majority class can post a
high accuracy while never catching the class you built it for.
Macro-F1 averages each class's F1 with equal weight, so that model
scores poorly, which is what you want your metric to tell you.
What this benchmark does and doesn't tell you
It's a controlled comparison: one base model, one GPU, seven
public datasets, four approaches. That makes it good evidence for the
relative ranking of approaches on short-text classification at small
model scale. It doesn't tell you how a 7B or 70B prompted model would
do (better than 0.6B, at many times the serving cost), and your own
data will be messier than a curated public benchmark. Your taxonomy
will have more edge cases, and your labels more disagreement. The
right move is to rerun the same comparison on a few hundred of your
own examples. The setup above is small enough to do it in a day.
FAQ
What is a financial text classifier?
A financial text classifier is a model that assigns each piece of financial text, such as a news sentence, a client message, or an analyst note, to one label from a fixed set: positive, neutral, or negative sentiment; one of dozens of banking intents; spam or not spam. Common uses include sentiment scoring on news and filings, routing client requests, and flagging suspicious communications for surveillance review.
Can I just prompt an LLM to classify financial text?
Not reliably with a small model. In Numerata's benchmark, a prompted 0.6B base model scored 29.9 F1 on three-way Financial PhraseBank sentiment, roughly chance, and 11.6 on 77-way Banking77 intent. LoRA fine-tuning the identical model raised those to 82.1 and 91.2. Larger hosted models prompt better, but cost more per call and send your text outside your environment.
How many labeled examples do I need to train a text classifier?
Fewer than most teams expect when fine-tuning a pretrained model with LoRA. On TREC-QC, 120 labeled examples got the LoRA fine-tune to 83.6 macro-F1, 35.5 points ahead of full fine-tuning on the same data. It reached 95.5 at 480 examples; full fine-tuning needed the full 5,452-example training set to get there, so the LoRA run used 91% fewer labels.
Is LoRA fine-tuning as accurate as full fine-tuning for classification?
In this benchmark it was more accurate on six of seven tasks and within 0.8 points on the seventh. Averaged across all seven tasks, the LoRA fine-tune scored 92.2 F1 against 89.3 for full fine-tuning. The gap was largest where training data is small relative to the label space, such as Banking77, where LoRA scored 91.2 against 85.4.
Is TF-IDF still a useful baseline for financial text classification?
Yes. TF-IDF with a linear classifier scored 94.0 on SpamAssassin, beating full fine-tuning, and 86.6 on Banking77, beating full fine-tuning there too. It is weakest where meaning depends on phrasing rather than vocabulary, such as financial sentiment, where it scored 66.5. Train it first: it takes minutes and tells you how much a model has to beat.
How fast is a fine-tuned small-model classifier?
The merged LoRA classifier in this benchmark ran at about 30 ms P50 latency at batch size one on a single GH200, compared with 63 to 174 ms for prompting the same base model. Merging the adapter into the base weights means there is no adapter overhead at serving time.
Why use macro-F1 instead of accuracy for financial classifiers?
Financial label distributions are usually imbalanced: most messages are not spam, most news sentences are neutral. Accuracy rewards a model for predicting the majority class. Macro-F1 averages the F1 score of every class with equal weight, so a model that never predicts the rare class, often the one you care about, scores poorly even when its accuracy looks high.
Where this fits at Numerata
The benchmark was run on Numerata's own stack: the LoRA training
on P95, the training layer, and the merged
classifiers served on NinetyFive, the inference
runtime, all on one GH200 inside the environment. Your labeled data,
your training runs, and the finished model stay on infrastructure you
control: private cloud, on-prem, or fully air-gapped.