# Llama-2-70B-chat (RAGTruth, 4-bit) Benchmark Scores & Performance

> Llama-2-70B-chat (RAGTruth, 4-bit) has published results in 2 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-ragtruth-llama-2-70b-chat-ragtruth-4-bit

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | Meta / TheBloke |
| Source Type | Open Weight |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Quantized checkpoint identified in RAGTruth](https://huggingface.co/TheBloke/Llama-2-70B-Chat-AWQ) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Llama-2-70B-chat (RAGTruth, 4-bit)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-ragtruth-llama-2-70b-chat-ragtruth-4-bit) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-ragtruth-llama-2-70b-chat-ragtruth-4-bit&format=csv)

### Generator hallucination counts and density

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.

Lower density means fewer hallucination spans per 100 response words. Counts are hallucinating responses and spans, not total requests. Llama-2-70B-chat uses TheBloke’s 4-bit AWQ checkpoint.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/) · [Full RAGTruth results](/benchmarks/ragtruth)

| Generator | QA hallucinating responses | QA hallucination spans | QA density | Data-to-text hallucinating responses | Data-to-text hallucination spans | Data-to-text density | Summarization hallucinating responses | Summarization hallucination spans | Summarization density | Overall hallucinating responses | Overall hallucination spans |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Llama-2-70B-chat † | 320 | 529 | 0.40 | 863 | 1834 | 1.15 | 212 | 245 | 0.26 | 1395 | 2608 |

### Hallucination annotation subtypes

The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.

Proportions are 0–1 ratios. Blank source cells are unreported, not zero.

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/). Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://aclanthology.org/2024.acl-long.585/) · [Full RAGTruth results](/benchmarks/ragtruth)

| Task | Generator | Hallucination spans | Implicit-true spans | Implicit-true proportion | Due-to-null spans | Due-to-null proportion |
| --- | --- | --- | --- | --- | --- | --- |
| Question Answering | Llama-2-70B-chat | 529 | 168 | 0.318 |  |  |
| Data-to-text Writing | Llama-2-70B-chat | 1834 | 274 | 0.149 | 272 | 0.148 |
| Summarization | Llama-2-70B-chat | 245 | 27 | 0.110 |  |  |
