# LMvLM (gpt-4-turbo) (RAGTruth detector) Benchmark Scores & Performance

> LMvLM (gpt-4-turbo) (RAGTruth detector) has published results in 1 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-ragtruth-lmvlm-gpt-4-turbo-detector

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | RAGTruth authors |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://aclanthology.org/2024.acl-long.585/) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: LMvLM (gpt-4-turbo) (RAGTruth detector)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-ragtruth-lmvlm-gpt-4-turbo-detector) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-ragtruth-lmvlm-gpt-4-turbo-detector&format=csv)

### Response-level hallucination detection

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.



RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/) · [Full RAGTruth results](/benchmarks/ragtruth)

| Detector | QA precision (%) | QA recall (%) | QA F1 (%) | Data-to-text precision (%) | Data-to-text recall (%) | Data-to-text F1 (%) | Summarization precision (%) | Summarization recall (%) | Summarization F1 (%) | Overall precision (%) | Overall recall (%) | Overall F1 (%) |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| LMvLM (gpt-4-turbo) | 18.7 | 76.9 | 30.1 | 68.0 | 76.7 | 72.1 | 23.3 | 81.9 | 36.2 | 36.2 | 77.8 | 49.4 |

## Other RAGTruth authors Models

- [Finetuned Llama-2-13B (RAGTruth detector)](/models/native-ragtruth-finetuned-llama-2-13b-detector) - Score: not computed
- [GPT-3.5-turbo-0613 (RAGTruth generator)](/models/native-ragtruth-gpt-3-5-turbo-0613-ragtruth-generator) - Score: not computed
- [GPT-4-0613 (RAGTruth generator)](/models/native-ragtruth-gpt-4-0613-ragtruth-generator) - Score: not computed
- [Llama-2-13B-chat (RAGTruth generator)](/models/native-ragtruth-llama-2-13b-chat-ragtruth-generator) - Score: not computed
- [Llama-2-7B-chat (RAGTruth generator)](/models/native-ragtruth-llama-2-7b-chat-ragtruth-generator) - Score: not computed
- [Mistral-7B-Instruct (RAGTruth generator)](/models/native-ragtruth-mistral-7b-instruct-ragtruth-generator) - Score: not computed
- [Prompt (gpt-3.5-turbo) (RAGTruth detector)](/models/native-ragtruth-prompt-gpt-3-5-turbo-detector) - Score: not computed
- [Prompt (gpt-4-turbo) (RAGTruth detector)](/models/native-ragtruth-prompt-gpt-4-turbo-detector) - Score: not computed
- [SelfCheckGPT (gpt-3.5-turbo) (RAGTruth detector)](/models/native-ragtruth-selfcheckgpt-gpt-3-5-turbo-detector) - Score: not computed
