Skip to main content
BenchLM

RAGTruth

We show this table for reference; we do not rank on it.

RAGTruth annotates unsupported or contradictory content in answers generated with retrieved evidence. It supports evaluation of hallucination detectors across question answering, summarization, and data-to-text tasks.

Original benchmark results

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.

8 source tables · 11 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)

Generator hallucination counts and density

Lower density means fewer hallucination spans per 100 response words. Counts are hallucinating responses and spans, not total requests. Llama-2-70B-chat uses TheBloke’s 4-bit AWQ checkpoint.

Published source

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Generator hallucination counts and density. Published source metrics with original precision; unreported cells are not zero.
GeneratorQA hallucinating responsesQA hallucination spansQA densityData-to-text hallucinating responsesData-to-text hallucination spansData-to-text densitySummarization hallucinating responsesSummarization hallucination spansSummarization densityOverall hallucinating responsesOverall hallucination spans
GPT-3.5-turbo-061375890.122723840.1854600.05401533
GPT-4-061348510.062903540.2774800.08406485
Llama-2-7B-chat51010100.5988817751.274345170.5818323302
Llama-2-13B-chat3996540.4898328031.532953420.4116773799
Llama-2-70B-chat †3205290.4086318341.152122450.2613952608
Mistral-7B-Instruct3785940.5995821401.516178280.8619533562

Context length bins and hallucination density

The source includes token intervals beside each density; these are source summaries, not new model accuracy results.

Published source

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Context length bins and hallucination density. Published source metrics with original precision; unreported cells are not zero.
Context length binSummarization density (token interval)Data-to-text density (token interval)QA density (token interval)
10.29_{(176,368]}1.51_{(178,273]}0.50_{(131,187]}
20.36_{(368,587]}1.48_{(273,378]}0.51_{(187,288]}
30.44_{(587,1422]}1.49_{(378,731]}0.49_{(288,400]}

Response length bins and hallucination density

The source includes token intervals beside each density; these are source summaries, not new model accuracy results.

Published source

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Response length bins and hallucination density. Published source metrics with original precision; unreported cells are not zero.
Response length binSummarization density (token interval)Data-to-text density (token interval)QA density (token interval)
10.34_{(44,87]}1.20_{(93,131]}0.21_{(19,93]}
20.32_{(87,119]}1.59_{(131,175]}0.37_{(93,138]}
30.44_{(119,245]}1.69_{(175,258]}0.87_{(138,257]}

Response-level hallucination detection

Published source

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Response-level hallucination detection. Published source metrics with original precision; unreported cells are not zero.
DetectorQA precision (%)QA recall (%)QA F1 (%)Data-to-text precision (%)Data-to-text recall (%)Data-to-text F1 (%)Summarization precision (%)Summarization recall (%)Summarization F1 (%)Overall precision (%)Overall recall (%)Overall F1 (%)
Prompt (gpt-3.5-turbo)18.884.430.865.195.577.423.489.237.137.192.352.9
Prompt (gpt-4-turbo)33.290.645.664.3100.078.331.597.647.646.997.963.4
SelfCheckGPT (gpt-3.5-turbo)35.058.043.768.282.874.831.156.540.149.771.958.8
LMvLM (gpt-4-turbo)18.776.930.168.076.772.123.381.936.236.277.849.4
Finetuned Llama-2-13B61.676.368.285.491.088.164.054.959.176.980.778.7

Span-level hallucination detection

Published source

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Span-level hallucination detection. Published source metrics with original precision; unreported cells are not zero.
DetectorQA precision (%)QA recall (%)QA F1 (%)Data-to-text precision (%)Data-to-text recall (%)Data-to-text F1 (%)Summarization precision (%)Summarization recall (%)Summarization F1 (%)Overall precision (%)Overall recall (%)Overall F1 (%)
Prompt Baseline (gpt-3.5-turbo)7.925.112.18.745.114.66.133.710.37.835.312.8
Prompt Baseline (gpt-4-turbo)23.752.032.617.966.428.214.765.424.118.460.928.3
Finetuned Llama-2-13B55.860.858.256.550.753.552.430.838.855.650.252.7

Response-selection pipeline outcomes

Each row measures response selection over a generator pair, using the finetuned Llama-2-13B detector. Source daggers mark cases where no response met the no-hallucination criterion: 328 and 448 valid responses remain. This pipeline result is not a score for one model.

Published source

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Response-selection pipeline outcomes. Published source metrics with original precision; unreported cells are not zero.
Generator pair and baseline hallucination rates (%)Selection strategyValid responsesHallucination rate (%) and relative decrease
Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6)Random45052.4(-)
Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6)Select the response with fewer detected hallucination spans45041.1( ↓ 21.6% )
Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6)Select the response with no detected hallucination spans328 †19.3( ↓ 63.2% )
GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3)Random4509.8(-)
GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3)Select the response with fewer detected hallucination spans4505.6( ↓ 42.9% )
GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3)Select the response with no detected hallucination spans448 †4.8( ↓ 51.0% )

Hallucination annotation subtypes

Proportions are 0–1 ratios. Blank source cells are unreported, not zero.

Published source

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Hallucination annotation subtypes. Published source metrics with original precision; unreported cells are not zero.
TaskGeneratorHallucination spansImplicit-true spansImplicit-true proportionDue-to-null spansDue-to-null proportion
Question AnsweringGPT-3.5-turbo-061389330.371UnreportedUnreported
Question AnsweringGPT-4-061351150.294UnreportedUnreported
Question AnsweringLlama-2-7B-chat10102510.249UnreportedUnreported
Question AnsweringLlama-2-13B-chat6542150.329UnreportedUnreported
Question AnsweringLlama-2-70B-chat5291680.318UnreportedUnreported
Question AnsweringMistral-7B-Instruct5941640.276UnreportedUnreported
Data-to-text WritingGPT-3.5-turbo-0613384520.135690.180
Data-to-text WritingGPT-4-0613354240.0682090.590
Data-to-text WritingLlama-2-7B-chat17751950.1102300.130
Data-to-text WritingLlama-2-13B-chat28032600.094390.157
Data-to-text WritingLlama-2-70B-chat18342740.1492720.148
Data-to-text WritingMistral-7B-Instruct21401020.0484230.198
SummarizationGPT-3.5-turbo-061360140.233UnreportedUnreported
SummarizationGPT-4-061380100.125UnreportedUnreported
SummarizationLlama-2-7B-chat517440.085UnreportedUnreported
SummarizationLlama-2-13B-chat342280.082UnreportedUnreported
SummarizationLlama-2-70B-chat245270.110UnreportedUnreported
SummarizationMistral-7B-Instruct828520.063UnreportedUnreported
OverallUnreported1428919280.13516420.115

Detector recall by hallucination type (Figure 4)

Exact percentage labels printed above the bars in Figure 4, page 10869. Recall measures character overlap with labeled hallucination spans. Chart coordinates are not used to estimate scores.

Published source

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Detector recall by hallucination type (Figure 4). Published source metrics with original precision; unreported cells are not zero.
DetectorSubtle conflict recall (%)Evident conflict recall (%)Subtle baseless information recall (%)Evident baseless information recall (%)
Prompt (gpt-3.5-turbo)11.535.325.436.2
Prompt (gpt-4-turbo)63.466.349.860.4
Finetuned Llama-2-13B2.538.352.955.8

Original benchmark results and configurations

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.

Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.

Snapshot

8 source tables11 configurationsDisplay only

About RAGTruth

Year

2024

Tasks

Detecting unsupported claims in retrieved-context answers

Format

Decision classification

Difficulty

Depends on the evaluated split and label mapping

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.

Freshness and provenance

Version

RAGTruth 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Which RAGTruth results are included?

This page includes 8 numeric result and configuration tables from RAGTruth’s original paper and available benchmark-owner updates, covering 11 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

Can I compare these RAGTruth scores with the Perplexity panel?

Compare RAGTruth scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

Where can I download the RAGTruth results?

The results download on this page provides every imported RAGTruth table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.

Last updated: October 1, 2026 source review · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.