# W-Loughran (FinancialPhraseBank) Benchmark Scores & Performance

> W-Loughran (FinancialPhraseBank) has published results in 2 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-financialphrasebank-w-loughran

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | Malo et al. |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://arxiv.org/abs/1307.5336) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: W-Loughran (FinancialPhraseBank)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-financialphrasebank-w-loughran) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-financialphrasebank-w-loughran&format=csv)

### Class metrics at 100% and >75% agreement

The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.

All metrics are 0–1 ratios. SVM-MPQA is the paper’s baseline marked with footnote a.

Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts — Pekka Malo and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/1307.5336) · [Full FinancialPhraseBank results](/benchmarks/financialphrasebank)

| Agreement | Class | Metric | W-Loughran |
| --- | --- | --- | --- |
| 100% | Positive | Accuracy | 0.755 |
| 100% | Positive | Recall | 0.125 |
| 100% | Positive | Precision | 0.563 |
| 100% | Positive | F1-score | 0.204 |
| 100% | Neutral | Accuracy | 0.625 |
| 100% | Neutral | Recall | 0.914 |
| 100% | Neutral | Precision | 0.635 |
| 100% | Neutral | F1-score | 0.750 |
| 100% | Negative | Accuracy | 0.849 |
| 100% | Negative | Recall | 0.162 |
| 100% | Negative | Precision | 0.360 |
| 100% | Negative | F1-score | 0.223 |
| >75% | Positive | Accuracy | 0.758 |
| >75% | Positive | Recall | 0.166 |
| >75% | Positive | Precision | 0.602 |
| >75% | Positive | F1-score | 0.260 |
| >75% | Neutral | Accuracy | 0.636 |
| >75% | Neutral | Recall | 0.904 |
| >75% | Neutral | Precision | 0.649 |
| >75% | Neutral | F1-score | 0.756 |
| >75% | Negative | Accuracy | 0.863 |
| >75% | Negative | Recall | 0.195 |
| >75% | Negative | Precision | 0.378 |
| >75% | Negative | F1-score | 0.257 |

### Class metrics at >66% and >50% agreement

The original study reports ten-fold cross-validation at four annotator-agreement thresholds. Its class-specific accuracy, precision, recall, and F1 are ratios from 0 to 1. These are not overall sentiment accuracy or Perplexity’s sampled panel. No unified modern leaderboard is published in the reviewed benchmark-owner sources.

All metrics are 0–1 ratios. SVM-MPQA is the paper’s baseline marked with footnote a.

Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts — Pekka Malo and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/1307.5336) · [Full FinancialPhraseBank results](/benchmarks/financialphrasebank)

| Agreement | Class | Metric | W-Loughran |
| --- | --- | --- | --- |
| >66% | Positive | Accuracy | 0.741 |
| >66% | Positive | Recall | 0.182 |
| >66% | Positive | Precision | 0.606 |
| >66% | Positive | F1-score | 0.279 |
| >66% | Neutral | Accuracy | 0.622 |
| >66% | Neutral | Recall | 0.894 |
| >66% | Neutral | Precision | 0.631 |
| >66% | Neutral | F1-score | 0.740 |
| >66% | Negative | Accuracy | 0.865 |
| >66% | Negative | Recall | 0.214 |
| >66% | Negative | Precision | 0.399 |
| >66% | Negative | F1-score | 0.278 |
| >50% | Positive | Accuracy | 0.732 |
| >50% | Positive | Recall | 0.187 |
| >50% | Positive | Precision | 0.573 |
| >50% | Positive | F1-score | 0.282 |
| >50% | Neutral | Accuracy | 0.613 |
| >50% | Neutral | Recall | 0.882 |
| >50% | Neutral | Precision | 0.623 |
| >50% | Neutral | F1-score | 0.730 |
| >50% | Negative | Accuracy | 0.863 |
| >50% | Negative | Recall | 0.222 |
| >50% | Negative | Precision | 0.407 |
| >50% | Negative | F1-score | 0.287 |

## Other Malo et al. Models

- [LPS (FinancialPhraseBank)](/models/native-financialphrasebank-lps) - Score: not computed
- [R-LPS (FinancialPhraseBank)](/models/native-financialphrasebank-r-lps) - Score: not computed
- [SVM-MPQA (FinancialPhraseBank)](/models/native-financialphrasebank-svm-mpqa) - Score: not computed
- [W-MPQA (FinancialPhraseBank)](/models/native-financialphrasebank-w-mpqa) - Score: not computed
