Skip to main content
BenchLM

FinancialPhraseBank

We show this table for reference; we do not rank on it.

Classify financial news sentences as positive, negative, or neutral. FinancialPhraseBank tests whether a model interprets financial context rather than treating every favorable-sounding word as positive.

Original benchmark results

The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.

2 source tables · 5 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)

Class metrics at 100% and >75% agreement

All metrics are 0–1 ratios. SVM-MPQA is the paper’s baseline marked with footnote a.

Published source

Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts — Pekka Malo and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Class metrics at 100% and >75% agreement. Published source metrics with original precision; unreported cells are not zero.
AgreementClassMetricW-MPQAW-LoughranSVM-MPQAR-LPSLPS
100%PositiveAccuracy0.6590.7550.7460.8580.869
100%PositiveRecall0.5190.1250.0600.6980.737
100%PositivePrecision0.3740.5630.4660.7280.742
100%PositiveF1-score0.4350.2040.1060.7130.739
100%NeutralAccuracy0.5860.6250.6520.8510.828
100%NeutralRecall0.5810.9140.9630.8870.868
100%NeutralPrecision0.6940.6350.6450.8720.854
100%NeutralF1-score0.6320.7500.7730.8800.861
100%NegativeAccuracy0.8290.8490.8700.9470.951
100%NegativeRecall0.3700.1620.2050.7990.789
100%NegativePrecision0.3640.3600.5340.8010.839
100%NegativeF1-score0.3670.2230.2960.8000.813
>75%PositiveAccuracy0.6510.7580.7440.8260.836
>75%PositiveRecall0.5570.1660.0770.6140.658
>75%PositivePrecision0.3780.6020.5110.6770.690
>75%PositiveF1-score0.4510.2600.1330.6440.674
>75%NeutralAccuracy0.5710.6360.6570.7990.792
>75%NeutralRecall0.5530.9040.9730.8650.837
>75%NeutralPrecision0.6940.6490.6500.8210.830
>75%NeutralF1-score0.6160.7560.7790.8420.833
>75%NegativeAccuracy0.8410.8630.8860.9390.945
>75%NegativeRecall0.3670.1950.1620.7070.800
>75%NegativePrecision0.3540.3780.6300.7710.760
>75%NegativeF1-score0.3600.2570.2580.7380.780

Class metrics at >66% and >50% agreement

All metrics are 0–1 ratios. SVM-MPQA is the paper’s baseline marked with footnote a.

Published source

Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts — Pekka Malo and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Class metrics at >66% and >50% agreement. Published source metrics with original precision; unreported cells are not zero.
AgreementClassMetricW-MPQAW-LoughranSVM-MPQAR-LPSLPS
>66%PositiveAccuracy0.6420.7410.7220.7990.798
>66%PositiveRecall0.5720.1820.0760.5920.816
>66%PositivePrecision0.3990.6060.4940.6510.599
>66%PositiveF1-score0.4700.2790.1320.6200.691
>66%NeutralAccuracy0.5630.6220.6320.7610.753
>66%NeutralRecall0.5340.8940.9660.8300.705
>66%NeutralPrecision0.6720.6310.6260.7860.858
>66%NeutralF1-score0.5950.7400.7590.8070.774
>66%NegativeAccuracy0.8450.8650.8830.9310.937
>66%NegativeRecall0.3790.2140.1420.6810.768
>66%NegativePrecision0.3700.3990.5840.7310.729
>66%NegativeF1-score0.3750.2780.2280.7050.748
>50%PositiveAccuracy0.6300.7320.7160.7760.786
>50%PositiveRecall0.5680.1870.0680.5000.535
>50%PositivePrecision0.3910.5730.4650.6290.645
>50%PositiveF1-score0.4630.2820.1190.5570.585
>50%NeutralAccuracy0.5480.6130.6230.7290.735
>50%NeutralRecall0.5160.8820.9650.8360.809
>50%NeutralPrecision0.6500.6230.6160.7410.760
>50%NeutralF1-score0.5750.7300.7520.7860.784
>50%NegativeAccuracy0.8480.8630.8800.9220.933
>50%NegativeRecall0.3740.2220.1340.6130.772
>50%NegativePrecision0.3880.4070.5740.7210.716
>50%NegativeF1-score0.3810.2870.2170.6620.743

Original benchmark results and configurations

The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.

Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.

Snapshot

2 source tables5 configurationsDisplay only

About FinancialPhraseBank

Year

2013

Tasks

Financial sentiment classification

Format

Decision classification

Difficulty

Depends on the evaluated split and label mapping

The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.

Freshness and provenance

Version

FinancialPhraseBank 2013

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Which FinancialPhraseBank results are included?

This page includes 2 numeric result and configuration tables from FinancialPhraseBank’s original paper and available benchmark-owner updates, covering 5 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

Can I compare these FinancialPhraseBank scores with the Perplexity panel?

Compare FinancialPhraseBank scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

Where can I download the FinancialPhraseBank results?

The results download on this page provides every imported FinancialPhraseBank table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.

Last updated: October 1, 2026 source review · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.