# Prompt (gpt-3.5-turbo) (RAGTruth detector) Benchmark Scores & Performance

> Prompt (gpt-3.5-turbo) (RAGTruth detector) has published results in 3 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-ragtruth-prompt-gpt-3-5-turbo-detector

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | RAGTruth authors |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://aclanthology.org/2024.acl-long.585/) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Prompt (gpt-3.5-turbo) (RAGTruth detector)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-ragtruth-prompt-gpt-3-5-turbo-detector) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-ragtruth-prompt-gpt-3-5-turbo-detector&format=csv)

### Response-level hallucination detection

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.



RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/) · [Full RAGTruth results](/benchmarks/ragtruth)

| Detector | QA precision (%) | QA recall (%) | QA F1 (%) | Data-to-text precision (%) | Data-to-text recall (%) | Data-to-text F1 (%) | Summarization precision (%) | Summarization recall (%) | Summarization F1 (%) | Overall precision (%) | Overall recall (%) | Overall F1 (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Prompt (gpt-3.5-turbo) | 18.8 | 84.4 | 30.8 | 65.1 | 95.5 | 77.4 | 23.4 | 89.2 | 37.1 | 37.1 | 92.3 | 52.9 |

### Span-level hallucination detection

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.



RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/) · [Full RAGTruth results](/benchmarks/ragtruth)

| Detector | QA precision (%) | QA recall (%) | QA F1 (%) | Data-to-text precision (%) | Data-to-text recall (%) | Data-to-text F1 (%) | Summarization precision (%) | Summarization recall (%) | Summarization F1 (%) | Overall precision (%) | Overall recall (%) | Overall F1 (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Prompt Baseline (gpt-3.5-turbo) | 7.9 | 25.1 | 12.1 | 8.7 | 45.1 | 14.6 | 6.1 | 33.7 | 10.3 | 7.8 | 35.3 | 12.8 |

### Detector recall by hallucination type (Figure 4)

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.

Exact percentage labels printed above the bars in Figure 4, page 10869. Recall measures character overlap with labeled hallucination spans. Chart coordinates are not used to estimate scores.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585.pdf) · [Full RAGTruth results](/benchmarks/ragtruth)

| Detector | Subtle conflict recall (%) | Evident conflict recall (%) | Subtle baseless information recall (%) | Evident baseless information recall (%) |
| --- | --- | --- | --- | --- |
| Prompt (gpt-3.5-turbo) | 11.5 | 35.3 | 25.4 | 36.2 |

## Other RAGTruth authors Models

- [Finetuned Llama-2-13B (RAGTruth detector)](/models/native-ragtruth-finetuned-llama-2-13b-detector) - Score: not computed
- [GPT-3.5-turbo-0613 (RAGTruth generator)](/models/native-ragtruth-gpt-3-5-turbo-0613-ragtruth-generator) - Score: not computed
- [GPT-4-0613 (RAGTruth generator)](/models/native-ragtruth-gpt-4-0613-ragtruth-generator) - Score: not computed
- [Llama-2-13B-chat (RAGTruth generator)](/models/native-ragtruth-llama-2-13b-chat-ragtruth-generator) - Score: not computed
- [Llama-2-7B-chat (RAGTruth generator)](/models/native-ragtruth-llama-2-7b-chat-ragtruth-generator) - Score: not computed
- [LMvLM (gpt-4-turbo) (RAGTruth detector)](/models/native-ragtruth-lmvlm-gpt-4-turbo-detector) - Score: not computed
- [Mistral-7B-Instruct (RAGTruth generator)](/models/native-ragtruth-mistral-7b-instruct-ragtruth-generator) - Score: not computed
- [Prompt (gpt-4-turbo) (RAGTruth detector)](/models/native-ragtruth-prompt-gpt-4-turbo-detector) - Score: not computed
- [SelfCheckGPT (gpt-3.5-turbo) (RAGTruth detector)](/models/native-ragtruth-selfcheckgpt-gpt-3-5-turbo-detector) - Score: not computed
