# FinancialPhraseBank

> Classify financial news sentences as positive, negative, or neutral. FinancialPhraseBank tests whether a model interprets financial context rather than treating every favorable-sounding word as positive.

Canonical page: https://benchlm.ai/benchmarks/financialphrasebank

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026 source review

## About FinancialPhraseBank

- Year: 2013
- Tasks: Financial sentiment classification
- Format: Decision classification
- Difficulty: Depends on the evaluated split and label mapping
- Paper: [Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts](https://arxiv.org/abs/1307.5336)

The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.

FinancialPhraseBank is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Original benchmark results

The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.

[Download full results (JSON)](/api/data/decision-benchmarks?benchmark=financialPhraseBank)

### Class metrics at 100% and >75% agreement

All metrics are 0–1 ratios. SVM-MPQA is the paper’s baseline marked with footnote a.

Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts — Pekka Malo and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/1307.5336)

| Agreement | Class | Metric | [W-MPQA](/models/native-financialphrasebank-w-mpqa) | [W-Loughran](/models/native-financialphrasebank-w-loughran) | [SVM-MPQA](/models/native-financialphrasebank-svm-mpqa) | [R-LPS](/models/native-financialphrasebank-r-lps) | [LPS](/models/native-financialphrasebank-lps) |
| --- | --- | --- | --- | --- | --- | --- | --- |
| 100% | Positive | Accuracy | 0.659 | 0.755 | 0.746 | 0.858 | 0.869 |
| 100% | Positive | Recall | 0.519 | 0.125 | 0.060 | 0.698 | 0.737 |
| 100% | Positive | Precision | 0.374 | 0.563 | 0.466 | 0.728 | 0.742 |
| 100% | Positive | F1-score | 0.435 | 0.204 | 0.106 | 0.713 | 0.739 |
| 100% | Neutral | Accuracy | 0.586 | 0.625 | 0.652 | 0.851 | 0.828 |
| 100% | Neutral | Recall | 0.581 | 0.914 | 0.963 | 0.887 | 0.868 |
| 100% | Neutral | Precision | 0.694 | 0.635 | 0.645 | 0.872 | 0.854 |
| 100% | Neutral | F1-score | 0.632 | 0.750 | 0.773 | 0.880 | 0.861 |
| 100% | Negative | Accuracy | 0.829 | 0.849 | 0.870 | 0.947 | 0.951 |
| 100% | Negative | Recall | 0.370 | 0.162 | 0.205 | 0.799 | 0.789 |
| 100% | Negative | Precision | 0.364 | 0.360 | 0.534 | 0.801 | 0.839 |
| 100% | Negative | F1-score | 0.367 | 0.223 | 0.296 | 0.800 | 0.813 |
| >75% | Positive | Accuracy | 0.651 | 0.758 | 0.744 | 0.826 | 0.836 |
| >75% | Positive | Recall | 0.557 | 0.166 | 0.077 | 0.614 | 0.658 |
| >75% | Positive | Precision | 0.378 | 0.602 | 0.511 | 0.677 | 0.690 |
| >75% | Positive | F1-score | 0.451 | 0.260 | 0.133 | 0.644 | 0.674 |
| >75% | Neutral | Accuracy | 0.571 | 0.636 | 0.657 | 0.799 | 0.792 |
| >75% | Neutral | Recall | 0.553 | 0.904 | 0.973 | 0.865 | 0.837 |
| >75% | Neutral | Precision | 0.694 | 0.649 | 0.650 | 0.821 | 0.830 |
| >75% | Neutral | F1-score | 0.616 | 0.756 | 0.779 | 0.842 | 0.833 |
| >75% | Negative | Accuracy | 0.841 | 0.863 | 0.886 | 0.939 | 0.945 |
| >75% | Negative | Recall | 0.367 | 0.195 | 0.162 | 0.707 | 0.800 |
| >75% | Negative | Precision | 0.354 | 0.378 | 0.630 | 0.771 | 0.760 |
| >75% | Negative | F1-score | 0.360 | 0.257 | 0.258 | 0.738 | 0.780 |

### Class metrics at >66% and >50% agreement

All metrics are 0–1 ratios. SVM-MPQA is the paper’s baseline marked with footnote a.

Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts — Pekka Malo and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/1307.5336)

| Agreement | Class | Metric | [W-MPQA](/models/native-financialphrasebank-w-mpqa) | [W-Loughran](/models/native-financialphrasebank-w-loughran) | [SVM-MPQA](/models/native-financialphrasebank-svm-mpqa) | [R-LPS](/models/native-financialphrasebank-r-lps) | [LPS](/models/native-financialphrasebank-lps) |
| --- | --- | --- | --- | --- | --- | --- | --- |
| >66% | Positive | Accuracy | 0.642 | 0.741 | 0.722 | 0.799 | 0.798 |
| >66% | Positive | Recall | 0.572 | 0.182 | 0.076 | 0.592 | 0.816 |
| >66% | Positive | Precision | 0.399 | 0.606 | 0.494 | 0.651 | 0.599 |
| >66% | Positive | F1-score | 0.470 | 0.279 | 0.132 | 0.620 | 0.691 |
| >66% | Neutral | Accuracy | 0.563 | 0.622 | 0.632 | 0.761 | 0.753 |
| >66% | Neutral | Recall | 0.534 | 0.894 | 0.966 | 0.830 | 0.705 |
| >66% | Neutral | Precision | 0.672 | 0.631 | 0.626 | 0.786 | 0.858 |
| >66% | Neutral | F1-score | 0.595 | 0.740 | 0.759 | 0.807 | 0.774 |
| >66% | Negative | Accuracy | 0.845 | 0.865 | 0.883 | 0.931 | 0.937 |
| >66% | Negative | Recall | 0.379 | 0.214 | 0.142 | 0.681 | 0.768 |
| >66% | Negative | Precision | 0.370 | 0.399 | 0.584 | 0.731 | 0.729 |
| >66% | Negative | F1-score | 0.375 | 0.278 | 0.228 | 0.705 | 0.748 |
| >50% | Positive | Accuracy | 0.630 | 0.732 | 0.716 | 0.776 | 0.786 |
| >50% | Positive | Recall | 0.568 | 0.187 | 0.068 | 0.500 | 0.535 |
| >50% | Positive | Precision | 0.391 | 0.573 | 0.465 | 0.629 | 0.645 |
| >50% | Positive | F1-score | 0.463 | 0.282 | 0.119 | 0.557 | 0.585 |
| >50% | Neutral | Accuracy | 0.548 | 0.613 | 0.623 | 0.729 | 0.735 |
| >50% | Neutral | Recall | 0.516 | 0.882 | 0.965 | 0.836 | 0.809 |
| >50% | Neutral | Precision | 0.650 | 0.623 | 0.616 | 0.741 | 0.760 |
| >50% | Neutral | F1-score | 0.575 | 0.730 | 0.752 | 0.786 | 0.784 |
| >50% | Negative | Accuracy | 0.848 | 0.863 | 0.880 | 0.922 | 0.933 |
| >50% | Negative | Recall | 0.374 | 0.222 | 0.134 | 0.613 | 0.772 |
| >50% | Negative | Precision | 0.388 | 0.407 | 0.574 | 0.721 | 0.716 |
| >50% | Negative | F1-score | 0.381 | 0.287 | 0.217 | 0.662 | 0.743 |

## FAQ

### Which FinancialPhraseBank results are included?

This page includes 2 numeric result and configuration tables from FinancialPhraseBank’s original paper and available benchmark-owner updates, covering 5 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

### Can I compare these FinancialPhraseBank scores with the Perplexity panel?

Compare FinancialPhraseBank scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

### Where can I download the FinancialPhraseBank results?

The results download on this page provides every imported FinancialPhraseBank table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
