# RAGTruth

> RAGTruth annotates unsupported or contradictory content in answers generated with retrieved evidence. It supports evaluation of hallucination detectors across question answering, summarization, and data-to-text tasks.

Canonical page: https://benchlm.ai/benchmarks/ragtruth

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026 source review

## About RAGTruth

- Year: 2024
- Tasks: Detecting unsupported claims in retrieved-context answers
- Format: Decision classification
- Difficulty: Depends on the evaluated split and label mapping
- Paper: [RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models](https://github.com/ParticleMedia/RAGTruth)

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.

RAGTruth is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Original benchmark results

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.

[Download full results (JSON)](/api/data/decision-benchmarks?benchmark=ragTruth)

### Generator hallucination counts and density

Lower density means fewer hallucination spans per 100 response words. Counts are hallucinating responses and spans, not total requests. Llama-2-70B-chat uses TheBloke’s 4-bit AWQ checkpoint.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/)

| Generator | QA hallucinating responses | QA hallucination spans | QA density | Data-to-text hallucinating responses | Data-to-text hallucination spans | Data-to-text density | Summarization hallucinating responses | Summarization hallucination spans | Summarization density | Overall hallucinating responses | Overall hallucination spans |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [GPT-3.5-turbo-0613](/models/native-ragtruth-gpt-3-5-turbo-0613-ragtruth-generator) | 75 | 89 | 0.12 | 272 | 384 | 0.18 | 54 | 60 | 0.05 | 401 | 533 |
| [GPT-4-0613](/models/native-ragtruth-gpt-4-0613-ragtruth-generator) | 48 | 51 | 0.06 | 290 | 354 | 0.27 | 74 | 80 | 0.08 | 406 | 485 |
| [Llama-2-7B-chat](/models/native-ragtruth-llama-2-7b-chat-ragtruth-generator) | 510 | 1010 | 0.59 | 888 | 1775 | 1.27 | 434 | 517 | 0.58 | 1832 | 3302 |
| [Llama-2-13B-chat](/models/native-ragtruth-llama-2-13b-chat-ragtruth-generator) | 399 | 654 | 0.48 | 983 | 2803 | 1.53 | 295 | 342 | 0.41 | 1677 | 3799 |
| [Llama-2-70B-chat †](/models/native-ragtruth-llama-2-70b-chat-ragtruth-4-bit) | 320 | 529 | 0.40 | 863 | 1834 | 1.15 | 212 | 245 | 0.26 | 1395 | 2608 |
| [Mistral-7B-Instruct](/models/native-ragtruth-mistral-7b-instruct-ragtruth-generator) | 378 | 594 | 0.59 | 958 | 2140 | 1.51 | 617 | 828 | 0.86 | 1953 | 3562 |

### Context length bins and hallucination density

The source includes token intervals beside each density; these are source summaries, not new model accuracy results.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/)

| Context length bin | Summarization density (token interval) | Data-to-text density (token interval) | QA density (token interval) |
| --- | --- | --- | --- |
| 1 | 0.29_{(176,368]} | 1.51_{(178,273]} | 0.50_{(131,187]} |
| 2 | 0.36_{(368,587]} | 1.48_{(273,378]} | 0.51_{(187,288]} |
| 3 | 0.44_{(587,1422]} | 1.49_{(378,731]} | 0.49_{(288,400]} |

### Response length bins and hallucination density

The source includes token intervals beside each density; these are source summaries, not new model accuracy results.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/)

| Response length bin | Summarization density (token interval) | Data-to-text density (token interval) | QA density (token interval) |
| --- | --- | --- | --- |
| 1 | 0.34_{(44,87]} | 1.20_{(93,131]} | 0.21_{(19,93]} |
| 2 | 0.32_{(87,119]} | 1.59_{(131,175]} | 0.37_{(93,138]} |
| 3 | 0.44_{(119,245]} | 1.69_{(175,258]} | 0.87_{(138,257]} |

### Response-level hallucination detection



RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/)

| Detector | QA precision (%) | QA recall (%) | QA F1 (%) | Data-to-text precision (%) | Data-to-text recall (%) | Data-to-text F1 (%) | Summarization precision (%) | Summarization recall (%) | Summarization F1 (%) | Overall precision (%) | Overall recall (%) | Overall F1 (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Prompt (gpt-3.5-turbo)](/models/native-ragtruth-prompt-gpt-3-5-turbo-detector) | 18.8 | 84.4 | 30.8 | 65.1 | 95.5 | 77.4 | 23.4 | 89.2 | 37.1 | 37.1 | 92.3 | 52.9 |
| [Prompt (gpt-4-turbo)](/models/native-ragtruth-prompt-gpt-4-turbo-detector) | 33.2 | 90.6 | 45.6 | 64.3 | 100.0 | 78.3 | 31.5 | 97.6 | 47.6 | 46.9 | 97.9 | 63.4 |
| [SelfCheckGPT (gpt-3.5-turbo)](/models/native-ragtruth-selfcheckgpt-gpt-3-5-turbo-detector) | 35.0 | 58.0 | 43.7 | 68.2 | 82.8 | 74.8 | 31.1 | 56.5 | 40.1 | 49.7 | 71.9 | 58.8 |
| [LMvLM (gpt-4-turbo)](/models/native-ragtruth-lmvlm-gpt-4-turbo-detector) | 18.7 | 76.9 | 30.1 | 68.0 | 76.7 | 72.1 | 23.3 | 81.9 | 36.2 | 36.2 | 77.8 | 49.4 |
| [Finetuned Llama-2-13B](/models/native-ragtruth-finetuned-llama-2-13b-detector) | 61.6 | 76.3 | 68.2 | 85.4 | 91.0 | 88.1 | 64.0 | 54.9 | 59.1 | 76.9 | 80.7 | 78.7 |

### Span-level hallucination detection



RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/)

| Detector | QA precision (%) | QA recall (%) | QA F1 (%) | Data-to-text precision (%) | Data-to-text recall (%) | Data-to-text F1 (%) | Summarization precision (%) | Summarization recall (%) | Summarization F1 (%) | Overall precision (%) | Overall recall (%) | Overall F1 (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Prompt Baseline (gpt-3.5-turbo)](/models/native-ragtruth-prompt-gpt-3-5-turbo-detector) | 7.9 | 25.1 | 12.1 | 8.7 | 45.1 | 14.6 | 6.1 | 33.7 | 10.3 | 7.8 | 35.3 | 12.8 |
| [Prompt Baseline (gpt-4-turbo)](/models/native-ragtruth-prompt-gpt-4-turbo-detector) | 23.7 | 52.0 | 32.6 | 17.9 | 66.4 | 28.2 | 14.7 | 65.4 | 24.1 | 18.4 | 60.9 | 28.3 |
| [Finetuned Llama-2-13B](/models/native-ragtruth-finetuned-llama-2-13b-detector) | 55.8 | 60.8 | 58.2 | 56.5 | 50.7 | 53.5 | 52.4 | 30.8 | 38.8 | 55.6 | 50.2 | 52.7 |

### Response-selection pipeline outcomes

Each row measures response selection over a generator pair, using the finetuned Llama-2-13B detector. Source daggers mark cases where no response met the no-hallucination criterion: 328 and 448 valid responses remain. This pipeline result is not a score for one model.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/)

| Generator pair and baseline hallucination rates (%) | Selection strategy | Valid responses | Hallucination rate (%) and relative decrease |
| --- | --- | --- | --- |
| Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6) | Random | 450 | 52.4(-) |
| Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6) | Select the response with fewer detected hallucination spans | 450 | 41.1( ↓ 21.6% ) |
| Llama-2-7B-chat (51.8) Mistral-7B-Instruct (57.6) | Select the response with no detected hallucination spans | 328 † | 19.3( ↓ 63.2% ) |
| GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3) | Random | 450 | 9.8(-) |
| GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3) | Select the response with fewer detected hallucination spans | 450 | 5.6( ↓ 42.9% ) |
| GPT-3.5-Turbo-0613 (10.9) GPT-4-0613 (9.3) | Select the response with no detected hallucination spans | 448 † | 4.8( ↓ 51.0% ) |

### Hallucination annotation subtypes

Proportions are 0–1 ratios. Blank source cells are unreported, not zero.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/)

| Task | Generator | Hallucination spans | Implicit-true spans | Implicit-true proportion | Due-to-null spans | Due-to-null proportion |
| --- | --- | --- | --- | --- | --- | --- |
| Question Answering | [GPT-3.5-turbo-0613](/models/native-ragtruth-gpt-3-5-turbo-0613-ragtruth-generator) | 89 | 33 | 0.371 |  |  |
| Question Answering | [GPT-4-0613](/models/native-ragtruth-gpt-4-0613-ragtruth-generator) | 51 | 15 | 0.294 |  |  |
| Question Answering | [Llama-2-7B-chat](/models/native-ragtruth-llama-2-7b-chat-ragtruth-generator) | 1010 | 251 | 0.249 |  |  |
| Question Answering | [Llama-2-13B-chat](/models/native-ragtruth-llama-2-13b-chat-ragtruth-generator) | 654 | 215 | 0.329 |  |  |
| Question Answering | [Llama-2-70B-chat](/models/native-ragtruth-llama-2-70b-chat-ragtruth-4-bit) | 529 | 168 | 0.318 |  |  |
| Question Answering | [Mistral-7B-Instruct](/models/native-ragtruth-mistral-7b-instruct-ragtruth-generator) | 594 | 164 | 0.276 |  |  |
| Data-to-text Writing | [GPT-3.5-turbo-0613](/models/native-ragtruth-gpt-3-5-turbo-0613-ragtruth-generator) | 384 | 52 | 0.135 | 69 | 0.180 |
| Data-to-text Writing | [GPT-4-0613](/models/native-ragtruth-gpt-4-0613-ragtruth-generator) | 354 | 24 | 0.068 | 209 | 0.590 |
| Data-to-text Writing | [Llama-2-7B-chat](/models/native-ragtruth-llama-2-7b-chat-ragtruth-generator) | 1775 | 195 | 0.110 | 230 | 0.130 |
| Data-to-text Writing | [Llama-2-13B-chat](/models/native-ragtruth-llama-2-13b-chat-ragtruth-generator) | 2803 | 260 | 0.09 | 439 | 0.157 |
| Data-to-text Writing | [Llama-2-70B-chat](/models/native-ragtruth-llama-2-70b-chat-ragtruth-4-bit) | 1834 | 274 | 0.149 | 272 | 0.148 |
| Data-to-text Writing | [Mistral-7B-Instruct](/models/native-ragtruth-mistral-7b-instruct-ragtruth-generator) | 2140 | 102 | 0.048 | 423 | 0.198 |
| Summarization | [GPT-3.5-turbo-0613](/models/native-ragtruth-gpt-3-5-turbo-0613-ragtruth-generator) | 60 | 14 | 0.233 |  |  |
| Summarization | [GPT-4-0613](/models/native-ragtruth-gpt-4-0613-ragtruth-generator) | 80 | 10 | 0.125 |  |  |
| Summarization | [Llama-2-7B-chat](/models/native-ragtruth-llama-2-7b-chat-ragtruth-generator) | 517 | 44 | 0.085 |  |  |
| Summarization | [Llama-2-13B-chat](/models/native-ragtruth-llama-2-13b-chat-ragtruth-generator) | 342 | 28 | 0.082 |  |  |
| Summarization | [Llama-2-70B-chat](/models/native-ragtruth-llama-2-70b-chat-ragtruth-4-bit) | 245 | 27 | 0.110 |  |  |
| Summarization | [Mistral-7B-Instruct](/models/native-ragtruth-mistral-7b-instruct-ragtruth-generator) | 828 | 52 | 0.063 |  |  |
| Overall |  | 14289 | 1928 | 0.135 | 1642 | 0.115 |

### Detector recall by hallucination type (Figure 4)

Exact percentage labels printed above the bars in Figure 4, page 10869. Recall measures character overlap with labeled hallucination spans. Chart coordinates are not used to estimate scores.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585.pdf)

| Detector | Subtle conflict recall (%) | Evident conflict recall (%) | Subtle baseless information recall (%) | Evident baseless information recall (%) |
| --- | --- | --- | --- | --- |
| [Prompt (gpt-3.5-turbo)](/models/native-ragtruth-prompt-gpt-3-5-turbo-detector) | 11.5 | 35.3 | 25.4 | 36.2 |
| [Prompt (gpt-4-turbo)](/models/native-ragtruth-prompt-gpt-4-turbo-detector) | 63.4 | 66.3 | 49.8 | 60.4 |
| [Finetuned Llama-2-13B](/models/native-ragtruth-finetuned-llama-2-13b-detector) | 2.5 | 38.3 | 52.9 | 55.8 |

## FAQ

### Which RAGTruth results are included?

This page includes 8 numeric result and configuration tables from RAGTruth’s original paper and available benchmark-owner updates, covering 11 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

### Can I compare these RAGTruth scores with the Perplexity panel?

Compare RAGTruth scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

### Where can I download the RAGTruth results?

The results download on this page provides every imported RAGTruth table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
