BLOG · EVALUATING

Financial text classifiers:
fine-tune, don't prompt.

A financial text classifier assigns each piece of financial text, a news sentence, a client chat message, a compliance-flagged email, to one label from a fixed set. The fastest way to build one today is to ask a general-purpose LLM for the label. The most accurate way, by a wide margin, is to fine-tune a small model on a few hundred labeled examples. In Numerata's benchmark across seven public classification tasks, a prompted 0.6B-parameter base model scored 29.9 F1 on Financial PhraseBank sentiment and 11.6 on 77-way Banking77 intent. A LoRA fine-tune of the identical model scored 82.1 and 91.2, beat both TF-IDF and full fine-tuning on six of seven tasks, and answered in about 30 ms.

This post walks through the four standard ways to build a text classifier, the full benchmark results, why each approach wins or loses where it does, and a reproducible recipe for building your own.

+52.2 PTS VS. PROMPTING

Financial PhraseBank sentiment, same 0.6B model, same hardware: 29.9 F1 prompted, 82.1 after LoRA fine-tuning.

6 OF 7 TASKS WON

Against TF-IDF and full fine-tuning. The seventh, PhraseBank, is within 0.8 points of full tuning.

91% FEWER LABELS

To reach 95.5 macro-F1 on TREC-QC: 480 labeled examples, against 5,452 for full fine-tuning.

What is a financial text classifier?

Classification is the workhorse of financial NLP. It's less glamorous than research agents or document Q&A, but it's where most of the volume is: every headline that needs a sentiment score, every client message that needs routing, every outbound communication that needs a surveillance flag. The output is one label from a closed set, so the task is easy to specify, easy to evaluate, and very easy to run millions of times a day. That last property is why model size and latency matter more here than almost anywhere else.

The seven benchmark tasks below are public datasets, but each one is a direct stand-in for a production workload on a trading desk or in a bank's operations stack.

BenchmarkClassesWhat it testsFinancial workload it stands in for
Financial PhraseBank3Sentiment of sentences from financial newsNews and filings sentiment signals, earnings-call tone
Banking7777Fine-grained intent of retail banking queriesClient service routing, chatbot triage, ticket tagging
SMS Spam2Spam vs. legitimate short messagesPhishing and fraud screening on SMS and chat
SpamAssassin2Spam vs. legitimate emailInbound email screening, communications surveillance pre-filter
AG News4Topic of a news articleNews feed routing by sector or asset class
Emotion6Emotion expressed in short textComplaint escalation, client-sentiment monitoring
TREC-QC6Type of answer a question expectsQuery routing for research assistants and internal search

Four ways to build a text classifier

Every approach in the benchmark turns text into a label. They differ in where the knowledge of your label set lives: in a prompt, in word counts, or in the model's weights.

  • Prompted base model (zero-shot). Describe the labels in a prompt and ask the model to answer with one. No training at all. The label set lives entirely in the context window, so the model has to read it on every call and then generate a label as text, which your code has to parse.
  • TF-IDF + linear classifier. The classic baseline. Represent each document as weighted word and n-gram counts, then fit logistic regression or a linear SVM. Trains in seconds on a CPU. It has no notion of word order or meaning beyond the n-grams it sees, which is a strength on keyword-driven tasks and a weakness on phrasing-driven ones.
  • Full fine-tuning. Update every weight in the pretrained model on your labeled data. Maximum flexibility, but with a few hundred examples and hundreds of millions of free parameters, it overfits easily and can overwrite useful pretrained knowledge.
  • LoRA fine-tuning, merged. Freeze the pretrained weights and train a small low-rank update alongside them. After training, fold the update back into the base weights so the deployed model is a single ordinary checkpoint. This is the approach labeled "Numerata" in the results.
Where the label set lives in each classification approach Four rows, one per approach. Prompting puts the label list in the context window and parses generated text. TF-IDF turns text into word counts and fits a linear model. Full fine-tuning updates all model weights. LoRA trains a small low-rank update and merges it into the weights before serving. INPUT WHERE THE LABEL SET LIVES OUTPUT TEXT PROMPT: "pick one of 77 labels..." RE-READ EVERY CALL · FROZEN WEIGHTS FREE TEXT PARSE, MAY FAIL TEXT N-GRAM COUNTS → LINEAR WEIGHTS NO WORD ORDER, NO PRETRAINING LABEL TEXT ALL WEIGHTS UPDATED HUNDREDS OF MILLIONS OF FREE PARAMETERS LABEL TEXT FROZEN W + LOW-RANK BA, MERGED SMALL UPDATE · ONE CHECKPOINT AT SERVE TIME LABEL
Prompting keeps the label set in the context window. The other three put it somewhere the model can learn it.

Benchmark setup

All four approaches were run against the same seven tasks. The three model-based approaches share one 0.6B-parameter base model, so every difference in the results comes from how the model was adapted, not from model size. Scores are F1 on each task's held-out test set, higher is better (macro-F1 for TREC-QC). Latency is P50 at batch size one. Every number was produced on a single NVIDIA GH200 inside a self-hosted environment, with no data or model leaving it.

Results: accuracy across seven tasks

TaskClassesPrompted baseTF-IDFFull tuneNumerata (LoRA)
SMS Spam20.090.294.697.1
AG News467.088.491.692.0
Financial PhraseBank329.966.582.982.1
Banking777711.686.685.491.2
SpamAssassin23.994.087.196.9
Emotion629.581.486.989.1
TREC-QC67.276.096.796.9
Mean of 721.383.389.392.2
F1 by task for four classification approaches on the same 0.6B model Dot plot of F1 scores on seven text-classification tasks. The prompted base model scores between 0 and 67. TF-IDF scores 66.5 to 94. Full fine-tuning scores 82.9 to 96.7. The LoRA fine-tune scores 82.1 to 97.1 and is highest on six of seven tasks; full tuning edges it on Financial PhraseBank, 82.9 to 82.1. PROMPTED BASE PROMPTED BASE TF-IDF TF-IDF FULL TUNE FULL TUNE NUMERATA (LORA, MERGED) NUMERATA (LORA, MERGED) 0 25 50 75 100 NUMERATA SMS Spam 2 CLASSES SMS Spam · Prompted base: 0.0 F1 SMS Spam · TF-IDF: 90.2 F1 SMS Spam · Full tune: 94.6 F1 SMS Spam · Numerata LoRA: 97.1 F1 97.1 AG News 4 CLASSES AG News · Prompted base: 67.0 F1 AG News · TF-IDF: 88.4 F1 AG News · Full tune: 91.6 F1 AG News · Numerata LoRA: 92.0 F1 92.0 Fin. PhraseBank 3 CLASSES Fin. PhraseBank · Prompted base: 29.9 F1 Fin. PhraseBank · TF-IDF: 66.5 F1 Fin. PhraseBank · Full tune: 82.9 F1 Fin. PhraseBank · Numerata LoRA: 82.1 F1 82.1 Banking77 77 CLASSES Banking77 · Prompted base: 11.6 F1 Banking77 · TF-IDF: 86.6 F1 Banking77 · Full tune: 85.4 F1 Banking77 · Numerata LoRA: 91.2 F1 91.2 SpamAssassin 2 CLASSES SpamAssassin · Prompted base: 3.9 F1 SpamAssassin · TF-IDF: 94.0 F1 SpamAssassin · Full tune: 87.1 F1 SpamAssassin · Numerata LoRA: 96.9 F1 96.9 Emotion 6 CLASSES Emotion · Prompted base: 29.5 F1 Emotion · TF-IDF: 81.4 F1 Emotion · Full tune: 86.9 F1 Emotion · Numerata LoRA: 89.1 F1 89.1 TREC-QC 6 CLASSES TREC-QC · Prompted base: 7.2 F1 TREC-QC · TF-IDF: 76.0 F1 TREC-QC · Full tune: 96.7 F1 TREC-QC · Numerata LoRA: 96.9 F1 96.9 F1, HIGHER IS BETTER
Same tasks as the table. Each row spans worst to best approach; hover a marker for its exact score.

Three patterns stand out. Prompting is not a weaker version of the other approaches, it's a different regime entirely, averaging 21.3 F1 against 83 to 92 for everything else. TF-IDF is a far stronger baseline than its reputation suggests, and beats full fine-tuning on two tasks. And the LoRA fine-tune is the only approach that is never meaningfully behind on any task.

Why prompting a small model fails at classification

The prompted scores are not a rounding error. On Financial PhraseBank, a three-way task, 29.9 F1 is roughly what you would get by guessing. On SMS Spam the prompted model scores 0.0, and on SpamAssassin 3.9: it essentially never catches the spam it was asked to find. Three mechanisms explain most of this.

  • The label set has to fit in working memory. For Banking77, the prompt has to enumerate 77 intents, many of them near-synonyms ("card not working" vs. "card swallowed" vs. "declined card payment"). A 0.6B model reading that list once per request has no stable internal representation of the boundaries between them. It scored 11.6. TF-IDF, which learned those boundaries from examples, scored 86.6.
  • Base-model priors don't match your taxonomy. Financial sentiment is not everyday sentiment. "Operating profit fell less than expected" is positive news. "The company will maintain its dividend" is usually neutral. A model that has never been shown your labeling convention defaults to the everyday reading.
  • Generated labels have to be parsed. A prompted model returns free text. Any response that isn't exactly a valid label, a hedge, an explanation, a label that isn't in the set, counts as a miss. Small models are especially prone to this, and on a binary task where the model drifts toward one answer, F1 on the positive class collapses toward zero.

Larger hosted models prompt considerably better than a 0.6B model does. But they cost far more per call, they're slower, and for most financial text they mean sending client communications or non-public information to a third party. Fine-tuning gets a small model to the accuracy you need, inside your own environment.

When TF-IDF is good enough, and when it isn't

TF-IDF scored 94.0 on SpamAssassin, beating full fine-tuning by 6.9 points, and 86.6 on Banking77, beating it by 1.2. Both are tasks where specific words carry most of the signal: spam has a vocabulary, and banking intents are usually named by their nouns ("top-up", "PIN", "exchange rate").

It's weakest where meaning depends on how words combine. On Financial PhraseBank it scored 66.5, more than 15 points behind both fine-tuned models, because "profit rose" and "loss rose" share most of their n-grams and mean opposite things. TREC-QC is similar: question type depends on structure ("how many" vs. "how did"), and TF-IDF trailed by about 21 points.

The practical rule: always train TF-IDF first. It takes minutes, it tells you how much a model has to beat to earn its serving cost, and on a keyword-driven task it may be all you need.

Why LoRA beat full fine-tuning

LoRA (low-rank adaptation) freezes each pretrained weight matrix W and learns an update ΔW = BA, where B and A are thin matrices of rank r, typically 8 to 64. For a 1,024 × 1,024 projection at rank 16, that's about 33,000 trainable parameters instead of about a million. Across a whole model, the trainable parameter count usually drops to around 1% of the total or less.

LoRA training and merging During training, a frozen weight matrix W is paired with two thin trainable matrices B and A whose product is a low-rank update. After training, the product BA is scaled and added into W to produce a single merged matrix W prime, which serves with no extra computation. TRAINING SERVING W FROZEN · d × k + B d × r A r × k TRAINABLE, r ≪ d, k MERGE W′ SAME SHAPE AS W W′ = W + (α / r) · BA NO ADAPTER AT INFERENCE, NO EXTRA LATENCY
Train a small update, fold it into the weights, serve one ordinary checkpoint.

That constraint is the point. With a few hundred or a few thousand labeled examples, full fine-tuning has far more freedom than the data can pin down, so it tends to memorize the training set and drift away from the general language knowledge that made the pretrained model useful. The low-rank update acts as a strong regularizer: the model can only move in a small number of directions, so it keeps its pretrained representations and learns the decision boundary on top of them.

The benchmark shows exactly where that matters. The gap between LoRA and full tuning is widest on SpamAssassin (+9.8) and Banking77 (+5.8), where training data is limited relative to the variety in the inputs or the size of the label space. It's narrowest on TREC-QC (+0.2) and AG News (+0.4), where full tuning had enough data to converge. On Financial PhraseBank full tuning edged ahead by 0.8, a tie for practical purposes.

Data efficiency: how many labels do you need?

For most financial teams, labeled data is the real bottleneck. Labels have to come from analysts or compliance staff whose time is expensive, and sometimes the text itself can't be sent to an outside labeling vendor. So the useful question is not "which approach is best with unlimited data" but "which approach is best with the 200 examples we can realistically label this month".

To answer it, the benchmark trained each approach on TREC-QC at five dataset sizes, from 60 labeled examples up to the full 5,452.

TREC-QC macro-F1 versus number of labeled training examples Learning curves for three approaches at 60, 120, 300, 480 and 5,452 labeled examples. At 120 labels the LoRA fine-tune scores 83.6 macro-F1 against 48.1 for full fine-tuning and 42.2 for TF-IDF. The LoRA fine-tune reaches 95.5 at 480 labels; full fine-tuning needs the entire 5,452-example training set to reach 96.7. 0 25 50 75 100 60 120 300 480 5,452 LABELED TRAINING EXAMPLES (NOT TO SCALE) TF-IDF · 60 labels: ~44.5 TF-IDF · 120 labels: 42.2 TF-IDF · 300 labels: ~61.0 TF-IDF · 480 labels: ~55.5 TF-IDF · 5,452 labels: 76.0 Full tune · 60 labels: ~22.0 Full tune · 120 labels: 48.1 Full tune · 300 labels: ~70.5 Full tune · 480 labels: ~80.5 Full tune · 5,452 labels: 96.7 Numerata LoRA · 60 labels: ~44.5 Numerata LoRA · 120 labels: 83.6 Numerata LoRA · 300 labels: ~89.0 Numerata LoRA · 480 labels: 95.5 Numerata LoRA · 5,452 labels: 96.9 83.6 48.1 42.2 95.5 @ 480 96.9 / 96.7 76.0 NUMERATA LORA NUMERATA LORA FULL TUNE FULL TUNE TF-IDF TF-IDF
TREC-QC macro-F1. Labeled values are exact; unlabeled points are read off the benchmark chart and approximate.

At 120 labeled examples, the LoRA fine-tune reached 83.6 macro-F1. Full fine-tuning on the same 120 examples scored about 48, and TF-IDF about 42: leads of 35.5 and 41.4 points. The LoRA run crossed 95.5 at 480 examples. Full fine-tuning needed the entire 5,452-example training set to reach a comparable score (96.7). The LoRA run got there with 91% fewer labels; put the other way, full fine-tuning needed more than 11 times the labeling for the same result.

That's the difference between a labeling project and a labeling afternoon. 120 examples is something a single analyst can produce in a few hours, and a dataset small enough to review line by line before it ever touches a model, which matters when compliance has to sign off on training data.

Latency: why the fine-tuned model is faster too

You might expect the more accurate model to be the slower one. Here it's the opposite.

P50 batch-1 latency, merged fine-tune versus prompting The merged LoRA classifier answers in about 30 milliseconds at P50, batch size one. Prompting the same base model takes between 63 and 174 milliseconds depending on the task. 0 50 100 150 200 MILLISECONDS, P50, BATCH 1 (LOWER IS BETTER) Merged fine-tune Merged fine-tune: ~30 ms P50 ~30 ms Prompted base RANGE ACROSS TASKS Prompted base: at least 63 ms P50 Prompted base: up to 174 ms P50, depending on task 63 to 174 ms
Same 0.6B base model, same GH200. The difference is entirely in what each request has to do.

The merged fine-tune runs at about 30 ms P50 at batch size one. Prompting the same base model takes 63 to 174 ms depending on the task, 2 to 6 times slower. Two things account for it:

  • Prompt length. A prompted classifier has to process the instructions and the full label list on every request before it reads the input. The more labels, the longer the prefill, which is why the prompted range is so wide across tasks. A fine-tuned classifier processes only the input text.
  • Output length. A prompted model generates a label as one or more tokens, often with preamble. A fine-tuned classifier produces a single prediction, one forward pass.

Merging matters for the same reason. Because the LoRA update is folded into the base weights before deployment, the served model has exactly the architecture and cost of the base model. There's no adapter lookup or extra matrix multiply per layer. At 30 ms a classifier can sit inline in a request path, scoring every headline as it arrives or every message before it's routed, instead of running as an overnight batch.

How to build a financial text classifier

The recipe below follows the benchmark's lessons. The code is a minimal, standard Hugging Face transformers + peft setup, one reasonable way to do it rather than the exact benchmark harness.

  1. Pin down the taxonomy. Write a one-line definition and two examples for every label, including the ambiguous cases. Most classifier failures in production are taxonomy disagreements, not model errors.
  2. Label a few hundred examples, stratified. Start with 100 to 500, making sure rare classes are represented. Hold out a separate test set of at least a few hundred that nobody trains on or tunes against.
  3. Train a TF-IDF baseline. It sets the bar the model has to clear.
  4. LoRA fine-tune a small model with a classification head.
  5. Evaluate with macro-F1 and a confusion matrix, not accuracy alone.
  6. Merge, serve, and monitor. Track the predicted label distribution over time. A shift is usually the first sign of input drift.

PYTHON · LORA CLASSIFIER

from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import LoraConfig, TaskType, get_peft_model

BASE = "path/to/small-base-model"            # any ~0.5-1B open-weight model
LABELS = ["negative", "neutral", "positive"]

tok = AutoTokenizer.from_pretrained(BASE)
if tok.pad_token is None:
    tok.pad_token = tok.eos_token

model = AutoModelForSequenceClassification.from_pretrained(
    BASE,
    num_labels=len(LABELS),
    id2label=dict(enumerate(LABELS)),
    label2id={l: i for i, l in enumerate(LABELS)},
)
model.config.pad_token_id = tok.pad_token_id

# SEQ_CLS keeps the new classification head fully trainable;
# only the low-rank updates are trained inside the transformer.
lora = LoraConfig(
    task_type=TaskType.SEQ_CLS,
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)
model = get_peft_model(model, lora)
model.print_trainable_parameters()

# ... train with transformers.Trainer on your labeled split ...

# Fold BA into W: one ordinary checkpoint, no adapter at serve time.
merged = model.merge_and_unload()
merged.save_pretrained("fin-sentiment-merged")
tok.save_pretrained("fin-sentiment-merged")

PYTHON · EVALUATION

from sklearn.metrics import classification_report, confusion_matrix, f1_score

# y_true, y_pred: label ids on the held-out test set.
# For a prompted baseline, map unparseable outputs to a "wrong" id
# rather than dropping them, or the comparison is flattering.
print("macro-F1:", f1_score(y_true, y_pred, average="macro"))
print(classification_report(y_true, y_pred, target_names=LABELS, digits=3))
print(confusion_matrix(y_true, y_pred))

A note on metrics. Financial label sets are almost always imbalanced: most messages aren't spam, most news sentences are neutral. A model that always predicts the majority class can post a high accuracy while never catching the class you built it for. Macro-F1 averages each class's F1 with equal weight, so that model scores poorly, which is what you want your metric to tell you.

What this benchmark does and doesn't tell you

It's a controlled comparison: one base model, one GPU, seven public datasets, four approaches. That makes it good evidence for the relative ranking of approaches on short-text classification at small model scale. It doesn't tell you how a 7B or 70B prompted model would do (better than 0.6B, at many times the serving cost), and your own data will be messier than a curated public benchmark. Your taxonomy will have more edge cases, and your labels more disagreement. The right move is to rerun the same comparison on a few hundred of your own examples. The setup above is small enough to do it in a day.

FAQ

What is a financial text classifier?

A financial text classifier is a model that assigns each piece of financial text, such as a news sentence, a client message, or an analyst note, to one label from a fixed set: positive, neutral, or negative sentiment; one of dozens of banking intents; spam or not spam. Common uses include sentiment scoring on news and filings, routing client requests, and flagging suspicious communications for surveillance review.

Can I just prompt an LLM to classify financial text?

Not reliably with a small model. In Numerata's benchmark, a prompted 0.6B base model scored 29.9 F1 on three-way Financial PhraseBank sentiment, roughly chance, and 11.6 on 77-way Banking77 intent. LoRA fine-tuning the identical model raised those to 82.1 and 91.2. Larger hosted models prompt better, but cost more per call and send your text outside your environment.

How many labeled examples do I need to train a text classifier?

Fewer than most teams expect when fine-tuning a pretrained model with LoRA. On TREC-QC, 120 labeled examples got the LoRA fine-tune to 83.6 macro-F1, 35.5 points ahead of full fine-tuning on the same data. It reached 95.5 at 480 examples; full fine-tuning needed the full 5,452-example training set to get there, so the LoRA run used 91% fewer labels.

Is LoRA fine-tuning as accurate as full fine-tuning for classification?

In this benchmark it was more accurate on six of seven tasks and within 0.8 points on the seventh. Averaged across all seven tasks, the LoRA fine-tune scored 92.2 F1 against 89.3 for full fine-tuning. The gap was largest where training data is small relative to the label space, such as Banking77, where LoRA scored 91.2 against 85.4.

Is TF-IDF still a useful baseline for financial text classification?

Yes. TF-IDF with a linear classifier scored 94.0 on SpamAssassin, beating full fine-tuning, and 86.6 on Banking77, beating full fine-tuning there too. It is weakest where meaning depends on phrasing rather than vocabulary, such as financial sentiment, where it scored 66.5. Train it first: it takes minutes and tells you how much a model has to beat.

How fast is a fine-tuned small-model classifier?

The merged LoRA classifier in this benchmark ran at about 30 ms P50 latency at batch size one on a single GH200, compared with 63 to 174 ms for prompting the same base model. Merging the adapter into the base weights means there is no adapter overhead at serving time.

Why use macro-F1 instead of accuracy for financial classifiers?

Financial label distributions are usually imbalanced: most messages are not spam, most news sentences are neutral. Accuracy rewards a model for predicting the majority class. Macro-F1 averages the F1 score of every class with equal weight, so a model that never predicts the rare class, often the one you care about, scores poorly even when its accuracy looks high.

Where this fits at Numerata

The benchmark was run on Numerata's own stack: the LoRA training on P95, the training layer, and the merged classifiers served on NinetyFive, the inference runtime, all on one GH200 inside the environment. Your labeled data, your training runs, and the finished model stay on infrastructure you control: private cloud, on-prem, or fully air-gapped.

For more on the techniques behind these results, see what fine-tuning is and when you need it, the hands-on guide to fine-tuning LLMs for finance, and what makes a small language model useful. If you're short on labels entirely, synthetic data for fine-tuning covers how to bootstrap a training set, and GPU sizing for inference covers what it takes to serve a classifier like this at volume.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog