# Llama-2-chat 70B (Belebele 89-language English instructions) Benchmark Scores & Performance

> Llama-2-chat 70B (Belebele 89-language English instructions) has published results in 1 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-belebele-llama-2-chat-70b-89-language-english-instructions

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | Belebele authors |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://arxiv.org/abs/2308.16884) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Llama-2-chat 70B (Belebele 89-language English instructions)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-belebele-llama-2-chat-70b-89-language-english-instructions) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-belebele-llama-2-chat-70b-89-language-english-instructions&format=csv)

### Translation and instruction-language comparisons

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

The first two rows cover 91 non-English languages; the instruction-language comparison covers 89. These subsets are distinct from the full 122-language result.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Model | Variant | Evaluation setting | Average accuracy (%) | Language share ≥50% accuracy (%) | Language share ≥70% accuracy (%) | English accuracy (%) |
| --- | --- | --- | --- | --- | --- | --- |
| Llama-2-chat | 70B | English Instructions | 44.9 | 37.1% | 3.4% | 78.8 |

## Other Belebele authors Models

- [BLOOMZ 7.1B (Belebele zero-shot)](/models/native-belebele-bloomz-7-1b-zero-shot) - Score: not computed
- [Falcon 40B (Belebele 5-shot)](/models/native-belebele-falcon-40b-5-shot) - Score: not computed
- [GPT3.5-turbo unk (Belebele zero-shot)](/models/native-belebele-gpt3-5-turbo-unk-zero-shot) - Score: not computed
- [InfoXLM large (550M) (Belebele English fine-tuning)](/models/native-belebele-infoxlm-large-550m-english-fine-tuning) - Score: not computed
- [InfoXLM large (550M) (Belebele Translate-Train-All)](/models/native-belebele-infoxlm-large-550m-translate-train-all) - Score: not computed
- [Llama 1 13B (Belebele 5-shot)](/models/native-belebele-llama-1-13b-5-shot) - Score: not computed
- [Llama 1 30B (Belebele 5-shot)](/models/native-belebele-llama-1-30b-5-shot) - Score: not computed
- [Llama 1 65B (Belebele 5-shot)](/models/native-belebele-llama-1-65b-5-shot) - Score: not computed
- [Llama 1 7B (Belebele 5-shot)](/models/native-belebele-llama-1-7b-5-shot) - Score: not computed
- [Llama 2 base 70B (Belebele 5-shot)](/models/native-belebele-llama-2-base-70b-5-shot) - Score: not computed
- [Llama-2-chat 70B (Belebele 89-language translated instructions)](/models/native-belebele-llama-2-chat-70b-89-language-translated-instructions) - Score: not computed
- [Llama-2-chat 70B (Belebele 91-language in-language)](/models/native-belebele-llama-2-chat-70b-91-language-in-language) - Score: not computed
- [Llama-2-chat 70B (Belebele 91-language Translate-Test)](/models/native-belebele-llama-2-chat-70b-91-language-translate-test) - Score: not computed
- [Llama-2-chat 70B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-70b-zero-shot) - Score: not computed
- [Llama-2-chat 7B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-7b-zero-shot) - Score: not computed
- [XLM-R large (550M) (Belebele English fine-tuning)](/models/native-belebele-xlm-r-large-550m-english-fine-tuning) - Score: not computed
- [XLM-R large (550M) (Belebele Translate-Train-All)](/models/native-belebele-xlm-r-large-550m-translate-train-all) - Score: not computed
- [XLM-V large (1.2B) (Belebele English fine-tuning)](/models/native-belebele-xlm-v-large-1-2b-english-fine-tuning) - Score: not computed
- [XLM-V large (1.2B) (Belebele Translate-Train-All)](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) - Score: not computed
