# Llama-2-chat 70B (Belebele 91-language Translate-Test) Benchmark Scores & Performance

> Llama-2-chat 70B (Belebele 91-language Translate-Test) has published results in 2 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-belebele-llama-2-chat-70b-91-language-translate-test

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | Belebele authors |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://arxiv.org/abs/2308.16884) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Llama-2-chat 70B (Belebele 91-language Translate-Test)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-belebele-llama-2-chat-70b-91-language-translate-test) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-belebele-llama-2-chat-70b-91-language-translate-test&format=csv)

### All translated-test language results

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

The 91 non-English-language subset also shows an English reference row. Its averages are not the 122-language averages. In-language appendix average is 44.0; the summary prints 44.1.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Language or summary metric | Llama-2-chat 70B (Belebele 91-language Translate-Test) |
| --- | --- |
| AVG | 57.1 |
| PCT Above 50 | 78.0% |
| PCT Above 70 | 2.2% |
| eng_Latn | 78.8 |
| fra_Latn | 70.6 |
| por_Latn | 69.9 |
| deu_Latn | 65.7 |
| ita_Latn | 66.1 |
| spa_Latn | 69.3 |
| cat_Latn | 67.0 |
| swe_Latn | 66.1 |
| rus_Cyrl | 67.3 |
| dan_Latn | 66.8 |
| nld_Latn | 67.2 |
| nob_Latn | 68.3 |
| ukr_Cyrl | 66.0 |
| ron_Latn | 67.0 |
| srp_Cyrl | 66.2 |
| bul_Cyrl | 67.7 |
| ces_Latn | 65.6 |
| hrv_Latn | 65.3 |
| fin_Latn | 61.1 |
| slv_Latn | 61.2 |
| zho_Hans | 71.2 |
| pol_Latn | 63.0 |
| ind_Latn | 64.8 |
| hun_Latn | 62.9 |
| vie_Latn | 59.4 |
| zho_Hant | 65.8 |
| slk_Latn | 66.2 |
| afr_Latn | 65.0 |
| jpn_Jpan | 54.8 |
| zsm_Latn | 67.0 |
| kor_Hang | 56.7 |
| mkd_Cyrl | 66.7 |
| ell_Grek | 67.6 |
| tgl_Latn | 62.2 |
| tur_Latn | 62.6 |
| arb_Arab | 60.7 |
| hin_Deva | 62.8 |
| pes_Arab | 59.6 |
| heb_Hebr | 62.0 |
| lvs_Latn | 60.9 |
| ceb_Latn | 62.6 |
| lit_Latn | 60.8 |
| hin_Latn | 52.7 |
| tha_Thai | 54.1 |
| isl_Latn | 58.1 |
| jav_Latn | 55.3 |
| urd_Arab | 59.4 |
| est_Latn | 59.4 |
| als_Latn | 63.1 |
| asm_Beng | 57.7 |
| swh_Latn | 57.8 |
| ben_Beng | 61.0 |
| sun_Latn | 50.8 |
| mar_Deva | 60.0 |
| kat_Geor | 57.7 |
| tam_Taml | 55.9 |
| urd_Latn | 43.0 |
| hat_Latn | 56.3 |
| azj_Latn | 55.6 |
| sin_Sinh | 57.7 |
| pan_Guru | 57.6 |
| npi_Deva | 62.0 |
| ckb_Arab | 51.3 |
| kaz_Cyrl | 53.2 |
| hau_Latn | 43.4 |
| hye_Armn | 58.0 |
| mya_Mymr | 46.6 |
| khk_Cyrl | 52.2 |
| guj_Gujr | 59.6 |
| lin_Latn | 40.3 |
| lug_Latn | 38.7 |
| ssw_Latn | 43.2 |
| khm_Khmr | 52.8 |
| plt_Latn | 46.7 |
| ben_Latn | 45.1 |
| som_Latn | 40.8 |
| pbt_Arab | 48.8 |
| zul_Latn | 44.4 |
| nso_Latn | 43.4 |
| tsn_Latn | 40.4 |
| yor_Latn | 37.7 |
| ibo_Latn | 35.3 |
| mal_Mlym | 63.0 |
| xho_Latn | 49.2 |
| fuv_Latn | 29.4 |
| gaz_Latn | 37.0 |
| ory_Orya | 57.8 |
| amh_Ethi | 50.4 |
| wol_Latn | 39.0 |
| tel_Telu | 54.3 |
| lao_Laoo | 47.4 |
| kan_Knda | 62.0 |

### Translation and instruction-language comparisons

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

The first two rows cover 91 non-English languages; the instruction-language comparison covers 89. These subsets are distinct from the full 122-language result.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Model | Variant | Evaluation setting | Average accuracy (%) | Language share ≥50% accuracy (%) | Language share ≥70% accuracy (%) | English accuracy (%) |
| --- | --- | --- | --- | --- | --- | --- |
| Llama-2-chat | 70B | Translate-Test | 57.1 | 78.0% | 2.2% | 78.8 |

## Other Belebele authors Models

- [BLOOMZ 7.1B (Belebele zero-shot)](/models/native-belebele-bloomz-7-1b-zero-shot) - Score: not computed
- [Falcon 40B (Belebele 5-shot)](/models/native-belebele-falcon-40b-5-shot) - Score: not computed
- [GPT3.5-turbo unk (Belebele zero-shot)](/models/native-belebele-gpt3-5-turbo-unk-zero-shot) - Score: not computed
- [InfoXLM large (550M) (Belebele English fine-tuning)](/models/native-belebele-infoxlm-large-550m-english-fine-tuning) - Score: not computed
- [InfoXLM large (550M) (Belebele Translate-Train-All)](/models/native-belebele-infoxlm-large-550m-translate-train-all) - Score: not computed
- [Llama 1 13B (Belebele 5-shot)](/models/native-belebele-llama-1-13b-5-shot) - Score: not computed
- [Llama 1 30B (Belebele 5-shot)](/models/native-belebele-llama-1-30b-5-shot) - Score: not computed
- [Llama 1 65B (Belebele 5-shot)](/models/native-belebele-llama-1-65b-5-shot) - Score: not computed
- [Llama 1 7B (Belebele 5-shot)](/models/native-belebele-llama-1-7b-5-shot) - Score: not computed
- [Llama 2 base 70B (Belebele 5-shot)](/models/native-belebele-llama-2-base-70b-5-shot) - Score: not computed
- [Llama-2-chat 70B (Belebele 89-language English instructions)](/models/native-belebele-llama-2-chat-70b-89-language-english-instructions) - Score: not computed
- [Llama-2-chat 70B (Belebele 89-language translated instructions)](/models/native-belebele-llama-2-chat-70b-89-language-translated-instructions) - Score: not computed
- [Llama-2-chat 70B (Belebele 91-language in-language)](/models/native-belebele-llama-2-chat-70b-91-language-in-language) - Score: not computed
- [Llama-2-chat 70B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-70b-zero-shot) - Score: not computed
- [Llama-2-chat 7B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-7b-zero-shot) - Score: not computed
- [XLM-R large (550M) (Belebele English fine-tuning)](/models/native-belebele-xlm-r-large-550m-english-fine-tuning) - Score: not computed
- [XLM-R large (550M) (Belebele Translate-Train-All)](/models/native-belebele-xlm-r-large-550m-translate-train-all) - Score: not computed
- [XLM-V large (1.2B) (Belebele English fine-tuning)](/models/native-belebele-xlm-v-large-1-2b-english-fine-tuning) - Score: not computed
- [XLM-V large (1.2B) (Belebele Translate-Train-All)](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) - Score: not computed
