FinancialPhraseBank
We show this table for reference; we do not rank on it.
Classify financial news sentences as positive, negative, or neutral. FinancialPhraseBank tests whether a model interprets financial context rather than treating every favorable-sounding word as positive.
Original benchmark results
The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.
2 source tables · 5 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)
Class metrics at 100% and >75% agreement
All metrics are 0–1 ratios. SVM-MPQA is the paper’s baseline marked with footnote a.
Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts — Pekka Malo and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Agreement | Class | Metric | W-MPQA | W-Loughran | SVM-MPQA | R-LPS | LPS |
|---|---|---|---|---|---|---|---|
| 100% | Positive | Accuracy | 0.659 | 0.755 | 0.746 | 0.858 | 0.869 |
| 100% | Positive | Recall | 0.519 | 0.125 | 0.060 | 0.698 | 0.737 |
| 100% | Positive | Precision | 0.374 | 0.563 | 0.466 | 0.728 | 0.742 |
| 100% | Positive | F1-score | 0.435 | 0.204 | 0.106 | 0.713 | 0.739 |
| 100% | Neutral | Accuracy | 0.586 | 0.625 | 0.652 | 0.851 | 0.828 |
| 100% | Neutral | Recall | 0.581 | 0.914 | 0.963 | 0.887 | 0.868 |
| 100% | Neutral | Precision | 0.694 | 0.635 | 0.645 | 0.872 | 0.854 |
| 100% | Neutral | F1-score | 0.632 | 0.750 | 0.773 | 0.880 | 0.861 |
| 100% | Negative | Accuracy | 0.829 | 0.849 | 0.870 | 0.947 | 0.951 |
| 100% | Negative | Recall | 0.370 | 0.162 | 0.205 | 0.799 | 0.789 |
| 100% | Negative | Precision | 0.364 | 0.360 | 0.534 | 0.801 | 0.839 |
| 100% | Negative | F1-score | 0.367 | 0.223 | 0.296 | 0.800 | 0.813 |
| >75% | Positive | Accuracy | 0.651 | 0.758 | 0.744 | 0.826 | 0.836 |
| >75% | Positive | Recall | 0.557 | 0.166 | 0.077 | 0.614 | 0.658 |
| >75% | Positive | Precision | 0.378 | 0.602 | 0.511 | 0.677 | 0.690 |
| >75% | Positive | F1-score | 0.451 | 0.260 | 0.133 | 0.644 | 0.674 |
| >75% | Neutral | Accuracy | 0.571 | 0.636 | 0.657 | 0.799 | 0.792 |
| >75% | Neutral | Recall | 0.553 | 0.904 | 0.973 | 0.865 | 0.837 |
| >75% | Neutral | Precision | 0.694 | 0.649 | 0.650 | 0.821 | 0.830 |
| >75% | Neutral | F1-score | 0.616 | 0.756 | 0.779 | 0.842 | 0.833 |
| >75% | Negative | Accuracy | 0.841 | 0.863 | 0.886 | 0.939 | 0.945 |
| >75% | Negative | Recall | 0.367 | 0.195 | 0.162 | 0.707 | 0.800 |
| >75% | Negative | Precision | 0.354 | 0.378 | 0.630 | 0.771 | 0.760 |
| >75% | Negative | F1-score | 0.360 | 0.257 | 0.258 | 0.738 | 0.780 |
Class metrics at >66% and >50% agreement
All metrics are 0–1 ratios. SVM-MPQA is the paper’s baseline marked with footnote a.
Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts — Pekka Malo and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Agreement | Class | Metric | W-MPQA | W-Loughran | SVM-MPQA | R-LPS | LPS |
|---|---|---|---|---|---|---|---|
| >66% | Positive | Accuracy | 0.642 | 0.741 | 0.722 | 0.799 | 0.798 |
| >66% | Positive | Recall | 0.572 | 0.182 | 0.076 | 0.592 | 0.816 |
| >66% | Positive | Precision | 0.399 | 0.606 | 0.494 | 0.651 | 0.599 |
| >66% | Positive | F1-score | 0.470 | 0.279 | 0.132 | 0.620 | 0.691 |
| >66% | Neutral | Accuracy | 0.563 | 0.622 | 0.632 | 0.761 | 0.753 |
| >66% | Neutral | Recall | 0.534 | 0.894 | 0.966 | 0.830 | 0.705 |
| >66% | Neutral | Precision | 0.672 | 0.631 | 0.626 | 0.786 | 0.858 |
| >66% | Neutral | F1-score | 0.595 | 0.740 | 0.759 | 0.807 | 0.774 |
| >66% | Negative | Accuracy | 0.845 | 0.865 | 0.883 | 0.931 | 0.937 |
| >66% | Negative | Recall | 0.379 | 0.214 | 0.142 | 0.681 | 0.768 |
| >66% | Negative | Precision | 0.370 | 0.399 | 0.584 | 0.731 | 0.729 |
| >66% | Negative | F1-score | 0.375 | 0.278 | 0.228 | 0.705 | 0.748 |
| >50% | Positive | Accuracy | 0.630 | 0.732 | 0.716 | 0.776 | 0.786 |
| >50% | Positive | Recall | 0.568 | 0.187 | 0.068 | 0.500 | 0.535 |
| >50% | Positive | Precision | 0.391 | 0.573 | 0.465 | 0.629 | 0.645 |
| >50% | Positive | F1-score | 0.463 | 0.282 | 0.119 | 0.557 | 0.585 |
| >50% | Neutral | Accuracy | 0.548 | 0.613 | 0.623 | 0.729 | 0.735 |
| >50% | Neutral | Recall | 0.516 | 0.882 | 0.965 | 0.836 | 0.809 |
| >50% | Neutral | Precision | 0.650 | 0.623 | 0.616 | 0.741 | 0.760 |
| >50% | Neutral | F1-score | 0.575 | 0.730 | 0.752 | 0.786 | 0.784 |
| >50% | Negative | Accuracy | 0.848 | 0.863 | 0.880 | 0.922 | 0.933 |
| >50% | Negative | Recall | 0.374 | 0.222 | 0.134 | 0.613 | 0.772 |
| >50% | Negative | Precision | 0.388 | 0.407 | 0.574 | 0.721 | 0.716 |
| >50% | Negative | F1-score | 0.381 | 0.287 | 0.217 | 0.662 | 0.743 |
Original benchmark results and configurations
The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.
Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.
Snapshot
About FinancialPhraseBank
Year
2013
Tasks
Financial sentiment classification
Format
Decision classification
Difficulty
Depends on the evaluated split and label mapping
The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.
Freshness and provenance
Version
FinancialPhraseBank 2013
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
Which FinancialPhraseBank results are included?
This page includes 2 numeric result and configuration tables from FinancialPhraseBank’s original paper and available benchmark-owner updates, covering 5 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.
Can I compare these FinancialPhraseBank scores with the Perplexity panel?
Compare FinancialPhraseBank scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.
Where can I download the FinancialPhraseBank results?
The results download on this page provides every imported FinancialPhraseBank table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.