Data as of October 1, 2026 · How the score is built
Llama-2-70B-chat (RAGTruth, 4-bit)
Decision readingLlama-2-70B-chat (RAGTruth, 4-bit) is tracked, but not publicly ranked yet. The original benchmark tables below include the published metrics for this evaluated configuration. Their distinct protocols remain outside general model rankings.
Llama-2-70B-chat (RAGTruth, 4-bit) will be repriced, updated or retired. Get each notice with its source and date. Follow model changes
Share or export
Original benchmark results
Every published metric for this configuration, with source precision and evaluation limits. Download results (JSON) · Numeric metrics (CSV)
Generator hallucination counts and density
The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.
Lower density means fewer hallucination spans per 100 response words. Counts are hallucinating responses and spans, not total requests. Llama-2-70B-chat uses TheBloke’s 4-bit AWQ checkpoint.
Published source · Full RAGTruth results
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Generator | QA hallucinating responses | QA hallucination spans | QA density | Data-to-text hallucinating responses | Data-to-text hallucination spans | Data-to-text density | Summarization hallucinating responses | Summarization hallucination spans | Summarization density | Overall hallucinating responses | Overall hallucination spans |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Llama-2-70B-chat † | 320 | 529 | 0.40 | 863 | 1834 | 1.15 | 212 | 245 | 0.26 | 1395 | 2608 |
Hallucination annotation subtypes
The original ACL report measures hallucination production, response-level detection, span-level detection, and response selection separately. Detection precision, recall, and F1 are percentages; hallucination density and counts are different measures. Generator and detector configurations keep separate profiles. The report does not identify every Mistral or detector API snapshot.
Proportions are 0–1 ratios. Blank source cells are unreported, not zero.
Published source · Full RAGTruth results
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models — Cheng Niu and colleagues. CC BY 4.0. Copyright 2024 Association for Computational Linguistics. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Task | Generator | Hallucination spans | Implicit-true spans | Implicit-true proportion | Due-to-null spans | Due-to-null proportion |
|---|---|---|---|---|---|---|
| Question Answering | Llama-2-70B-chat | 529 | 168 | 0.318 | Unreported | Unreported |
| Data-to-text Writing | Llama-2-70B-chat | 1834 | 274 | 0.149 | 272 | 0.148 |
| Summarization | Llama-2-70B-chat | 245 | 27 | 0.110 | Unreported | Unreported |
Lineage
The sequence follows explicit supersedes links. Each score is estimated for that model; a relative can inform a sparse estimate but never sets a floor, so a newer release can score below an earlier one. Scores and prices remain blank when the corresponding public row or first-party rate is unavailable.
Release date not sourced · you are here
Llama-2-70B-chat (RAGTruth, 4-bit)Not publicly ranked · Price not listed
Benchmark-system
Spec sheet
Each documented value carries its source. Missing fields stay visible as not sourced or not published, rather than disappearing from the page.
- API model ID
- Not publishedProvider pricing
- Context window
- Not published
- Maximum output
- Not sourced yet
- Knowledge cutoff
- Not sourced yet
- Input modalities
- Not sourced yet
- Output modalities
- Not sourced yet
- Parameters
- Not sourced yet
- Availability
- Not sourced yet
- Cloud regions
- Not tracked yet
- Lifecycle
- Tracked
- API capabilities
- Tool calling, structured outputs, and batch support are not tracked yet
- Prompt caching
- Not documented in the pricing recordProvider pricing
- Self-host
- Open weights available; hardware estimate not sourced
- Rate limits
- Not tracked yet
How to read this profile
The visual layer above carries the decisions. These notes preserve the model, ranking, coverage, and family context behind the numbers.
The original benchmark tables show this evaluated configuration’s published results. Training choices, prompts, response splits, and metrics remain attached to each table. The profile stays outside general model rankings.
Llama-2-70B-chat (RAGTruth, 4-bit) is a open weight model. No explicit reasoning mode is documented in this profile.
Evaluated system: Llama-2-70B-chat (RAGTruth, 4-bit). The linked benchmark-owner report supplies the full numeric results and protocol. This profile represents that evaluated configuration; it does not transfer results to a base model or another prompt. Context limit, current API tariff, and release date are not established by this evaluation.
The original benchmark results appear in the source tables above; they remain separate from weighted benchmark slots.
Last updated October 1, 2026. Runtime fields remain blank until a sourced snapshot exists.
Deployment options
Self-host and provider-specific paths stay separate from benchmark evidence so operating constraints are visible before a score becomes the whole decision.
Published weights are available, but BenchLM does not yet have a sourced parameter and VRAM profile for this exact model. Hardware cost estimates stay unavailable until that sizing record is complete.
Estimate VRAM from known parametersQuestions
How does Llama-2-70B-chat (RAGTruth, 4-bit) perform overall in AI benchmarks?
This configuration has original benchmark results in the source tables on this profile. Every published metric retains its evaluated configuration, source precision, and protocol. These results stay outside general model rankings; they do not produce a public overall score or transfer to a base model with different settings.
Is Llama-2-70B-chat (RAGTruth, 4-bit) open source?
Llama-2-70B-chat (RAGTruth, 4-bit) is an open-weight model from Meta / TheBloke. Its weights can be downloaded for local or hosted deployment, subject to the published license. Open weight does not automatically mean open source: training data and training code may remain private, and commercial restrictions can still apply.
Does Llama-2-70B-chat (RAGTruth, 4-bit) have full benchmark coverage on BenchLM?
This configuration has results in the original benchmark tables on this profile. Those tables preserve every imported source metric, but they do not fill the site’s general scoring slots. Other benchmarks remain unmeasured, and compatible configurations are required before comparing results across different reports.
What is the context window size of Llama-2-70B-chat (RAGTruth, 4-bit)?
Llama-2-70B-chat (RAGTruth, 4-bit)'s context window is not documented in a source tied to this exact model yet. The profile leaves the value unavailable instead of borrowing a limit from an earlier family member or an unverified route. Maximum output length remains a separate field.