RAGTruth
We show this table for reference; we do not rank on it.
RAGTruth annotates unsupported or contradictory content in answers generated with retrieved evidence. It supports evaluation of hallucination detectors across question answering, summarization, and data-to-text tasks.
Original benchmark results
The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.
8 source tables · 11 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)
Generator hallucination counts and density
Lower density means fewer hallucination spans per 100 response words. Counts are hallucinating responses and spans, not total requests. Llama-2-70B-chat uses TheBloke’s 4-bit AWQ checkpoint.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Generator | QA hallucinating responses | QA hallucination spans | QA density | Data-to-text hallucinating responses | Data-to-text hallucination spans | Data-to-text density | Summarization hallucinating responses | Summarization hallucination spans | Summarization density | Overall hallucinating responses | Overall hallucination spans |
|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-3.5-turbo-0613 | 75 | 89 | 0.12 | 272 | 384 | 0.18 | 54 | 60 | 0.05 | 401 | 533 |
| GPT-4-0613 | 48 | 51 | 0.06 | 290 | 354 | 0.27 | 74 | 80 | 0.08 | 406 | 485 |
| Llama-2-7B-chat | 510 | 1010 | 0.59 | 888 | 1775 | 1.27 | 434 | 517 | 0.58 | 1832 | 3302 |
| Llama-2-13B-chat | 399 | 654 | 0.48 | 983 | 2803 | 1.53 | 295 | 342 | 0.41 | 1677 | 3799 |
| Llama-2-70B-chat † | 320 | 529 | 0.40 | 863 | 1834 | 1.15 | 212 | 245 | 0.26 | 1395 | 2608 |
| Mistral-7B-Instruct | 378 | 594 | 0.59 | 958 | 2140 | 1.51 | 617 | 828 | 0.86 | 1953 | 3562 |
Context length bins and hallucination density
The source includes token intervals beside each density; these are source summaries, not new model accuracy results.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Context length bin | Summarization density (token interval) | Data-to-text density (token interval) | QA density (token interval) |
|---|---|---|---|
| 1 | 0.29_{(176,368]} | 1.51_{(178,273]} | 0.50_{(131,187]} |
| 2 | 0.36_{(368,587]} | 1.48_{(273,378]} | 0.51_{(187,288]} |
| 3 | 0.44_{(587,1422]} | 1.49_{(378,731]} | 0.49_{(288,400]} |
Response length bins and hallucination density
The source includes token intervals beside each density; these are source summaries, not new model accuracy results.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Response length bin | Summarization density (token interval) | Data-to-text density (token interval) | QA density (token interval) |
|---|---|---|---|
| 1 | 0.34_{(44,87]} | 1.20_{(93,131]} | 0.21_{(19,93]} |
| 2 | 0.32_{(87,119]} | 1.59_{(131,175]} | 0.37_{(93,138]} |
| 3 | 0.44_{(119,245]} | 1.69_{(175,258]} | 0.87_{(138,257]} |
Response-level hallucination detection
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Detector | QA precision (%) | QA recall (%) | QA F1 (%) | Data-to-text precision (%) | Data-to-text recall (%) | Data-to-text F1 (%) | Summarization precision (%) | Summarization recall (%) | Summarization F1 (%) | Overall precision (%) | Overall recall (%) | Overall F1 (%) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Prompt (gpt-3.5-turbo) | 18.8 | 84.4 | 30.8 | 65.1 | 95.5 | 77.4 | 23.4 | 89.2 | 37.1 | 37.1 | 92.3 | 52.9 |
| Prompt (gpt-4-turbo) | 33.2 | 90.6 | 45.6 | 64.3 | 100.0 | 78.3 | 31.5 | 97.6 | 47.6 | 46.9 | 97.9 | 63.4 |
| SelfCheckGPT (gpt-3.5-turbo) | 35.0 | 58.0 | 43.7 | 68.2 | 82.8 | 74.8 | 31.1 | 56.5 | 40.1 | 49.7 | 71.9 | 58.8 |
| LMvLM (gpt-4-turbo) | 18.7 | 76.9 | 30.1 | 68.0 | 76.7 | 72.1 | 23.3 | 81.9 | 36.2 | 36.2 | 77.8 | 49.4 |
| Finetuned Llama-2-13B | 61.6 | 76.3 | 68.2 | 85.4 | 91.0 | 88.1 | 64.0 | 54.9 | 59.1 | 76.9 | 80.7 | 78.7 |
Span-level hallucination detection
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Detector | QA precision (%) | QA recall (%) | QA F1 (%) | Data-to-text precision (%) | Data-to-text recall (%) | Data-to-text F1 (%) | Summarization precision (%) | Summarization recall (%) | Summarization F1 (%) | Overall precision (%) | Overall recall (%) | Overall F1 (%) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Prompt Baseline (gpt-3.5-turbo) | 7.9 | 25.1 | 12.1 | 8.7 | 45.1 | 14.6 | 6.1 | 33.7 | 10.3 | 7.8 | 35.3 | 12.8 |
| Prompt Baseline (gpt-4-turbo) | 23.7 | 52.0 | 32.6 | 17.9 | 66.4 | 28.2 | 14.7 | 65.4 | 24.1 | 18.4 | 60.9 | 28.3 |
| Finetuned Llama-2-13B | 55.8 | 60.8 | 58.2 | 56.5 | 50.7 | 53.5 | 52.4 | 30.8 | 38.8 | 55.6 | 50.2 | 52.7 |
Response-selection pipeline outcomes
Each row measures response selection over a generator pair, using the finetuned Llama-2-13B detector. Source daggers mark cases where no response met the no-hallucination criterion: 328 and 448 valid responses remain. This pipeline result is not a score for one model.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Generator pair and baseline hallucination rates (%) | Selection strategy | Valid responses | Hallucination rate (%) and relative decrease |
|---|---|---|---|
| Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6) | Random | 450 | 52.4(-) |
| Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6) | Select the response with fewer detected hallucination spans | 450 | 41.1( ↓ 21.6% ) |
| Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6) | Select the response with no detected hallucination spans | 328 † | 19.3( ↓ 63.2% ) |
| GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3) | Random | 450 | 9.8(-) |
| GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3) | Select the response with fewer detected hallucination spans | 450 | 5.6( ↓ 42.9% ) |
| GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3) | Select the response with no detected hallucination spans | 448 † | 4.8( ↓ 51.0% ) |
Hallucination annotation subtypes
Proportions are 0–1 ratios. Blank source cells are unreported, not zero.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Task | Generator | Hallucination spans | Implicit-true spans | Implicit-true proportion | Due-to-null spans | Due-to-null proportion |
|---|---|---|---|---|---|---|
| Question Answering | GPT-3.5-turbo-0613 | 89 | 33 | 0.371 | Unreported | Unreported |
| Question Answering | GPT-4-0613 | 51 | 15 | 0.294 | Unreported | Unreported |
| Question Answering | Llama-2-7B-chat | 1010 | 251 | 0.249 | Unreported | Unreported |
| Question Answering | Llama-2-13B-chat | 654 | 215 | 0.329 | Unreported | Unreported |
| Question Answering | Llama-2-70B-chat | 529 | 168 | 0.318 | Unreported | Unreported |
| Question Answering | Mistral-7B-Instruct | 594 | 164 | 0.276 | Unreported | Unreported |
| Data-to-text Writing | GPT-3.5-turbo-0613 | 384 | 52 | 0.135 | 69 | 0.180 |
| Data-to-text Writing | GPT-4-0613 | 354 | 24 | 0.068 | 209 | 0.590 |
| Data-to-text Writing | Llama-2-7B-chat | 1775 | 195 | 0.110 | 230 | 0.130 |
| Data-to-text Writing | Llama-2-13B-chat | 2803 | 260 | 0.09 | 439 | 0.157 |
| Data-to-text Writing | Llama-2-70B-chat | 1834 | 274 | 0.149 | 272 | 0.148 |
| Data-to-text Writing | Mistral-7B-Instruct | 2140 | 102 | 0.048 | 423 | 0.198 |
| Summarization | GPT-3.5-turbo-0613 | 60 | 14 | 0.233 | Unreported | Unreported |
| Summarization | GPT-4-0613 | 80 | 10 | 0.125 | Unreported | Unreported |
| Summarization | Llama-2-7B-chat | 517 | 44 | 0.085 | Unreported | Unreported |
| Summarization | Llama-2-13B-chat | 342 | 28 | 0.082 | Unreported | Unreported |
| Summarization | Llama-2-70B-chat | 245 | 27 | 0.110 | Unreported | Unreported |
| Summarization | Mistral-7B-Instruct | 828 | 52 | 0.063 | Unreported | Unreported |
| Overall | Unreported | 14289 | 1928 | 0.135 | 1642 | 0.115 |
Detector recall by hallucination type (Figure 4)
Exact percentage labels printed above the bars in Figure 4, page 10869. Recall measures character overlap with labeled hallucination spans. Chart coordinates are not used to estimate scores.
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Detector | Subtle conflict recall (%) | Evident conflict recall (%) | Subtle baseless information recall (%) | Evident baseless information recall (%) |
|---|---|---|---|---|
| Prompt (gpt-3.5-turbo) | 11.5 | 35.3 | 25.4 | 36.2 |
| Prompt (gpt-4-turbo) | 63.4 | 66.3 | 49.8 | 60.4 |
| Finetuned Llama-2-13B | 2.5 | 38.3 | 52.9 | 55.8 |
Original benchmark results and configurations
The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.
Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.
Snapshot
About RAGTruth
Year
2024
Tasks
Detecting unsupported claims in retrieved-context answers
Format
Decision classification
Difficulty
Depends on the evaluated split and label mapping
The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.
Freshness and provenance
Version
RAGTruth 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
Which RAGTruth results are included?
This page includes 8 numeric result and configuration tables from RAGTruth’s original paper and available benchmark-owner updates, covering 11 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.
Can I compare these RAGTruth scores with the Perplexity panel?
Compare RAGTruth scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.
Where can I download the RAGTruth results?
The results download on this page provides every imported RAGTruth table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.