# Llama-2-chat 70B (Belebele 91-language in-language) Benchmark Scores & Performance

> Llama-2-chat 70B (Belebele 91-language in-language) has published results in 2 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-belebele-llama-2-chat-70b-91-language-in-language

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | Belebele authors |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://arxiv.org/abs/2308.16884) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Llama-2-chat 70B (Belebele 91-language in-language)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-belebele-llama-2-chat-70b-91-language-in-language) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-belebele-llama-2-chat-70b-91-language-in-language&format=csv)

### All translated-test language results

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

The 91 non-English-language subset also shows an English reference row. Its averages are not the 122-language averages. In-language appendix average is 44.0; the summary prints 44.1.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Language or summary metric | Llama-2-chat 70B (Belebele 91-language in-language) |
| --- | --- |
| AVG | 44.0 |
| PCT Above 50 | 35.2% |
| PCT Above 70 | 2.2% |
| eng_Latn | 78.8 |
| fra_Latn | 72.2 |
| por_Latn | 70.2 |
| deu_Latn | 69.4 |
| ita_Latn | 68.6 |
| spa_Latn | 68.4 |
| cat_Latn | 68.2 |
| swe_Latn | 67.4 |
| rus_Cyrl | 67.0 |
| dan_Latn | 66.2 |
| nld_Latn | 66.2 |
| nob_Latn | 65.7 |
| ukr_Cyrl | 65.7 |
| ron_Latn | 65.6 |
| srp_Cyrl | 65.1 |
| bul_Cyrl | 65.0 |
| ces_Latn | 65.0 |
| hrv_Latn | 64.7 |
| fin_Latn | 62.7 |
| slv_Latn | 62.4 |
| zho_Hans | 62.4 |
| pol_Latn | 61.7 |
| ind_Latn | 61.3 |
| hun_Latn | 61.1 |
| vie_Latn | 59.6 |
| zho_Hant | 59.3 |
| slk_Latn | 58.8 |
| afr_Latn | 57.9 |
| jpn_Jpan | 56.6 |
| zsm_Latn | 56.4 |
| kor_Hang | 56.3 |
| mkd_Cyrl | 55.7 |
| ell_Grek | 50.7 |
| tgl_Latn | 49.6 |
| tur_Latn | 47.3 |
| arb_Arab | 42.3 |
| hin_Deva | 42.0 |
| pes_Arab | 41.8 |
| heb_Hebr | 41.4 |
| lvs_Latn | 41.0 |
| ceb_Latn | 40.6 |
| lit_Latn | 39.7 |
| hin_Latn | 39.2 |
| tha_Thai | 38.9 |
| isl_Latn | 38.0 |
| jav_Latn | 37.0 |
| urd_Arab | 37.0 |
| est_Latn | 36.6 |
| als_Latn | 36.0 |
| asm_Beng | 35.7 |
| swh_Latn | 35.1 |
| ben_Beng | 34.9 |
| sun_Latn | 34.9 |
| mar_Deva | 34.8 |
| kat_Geor | 34.6 |
| tam_Taml | 34.4 |
| urd_Latn | 34.1 |
| hat_Latn | 34.1 |
| azj_Latn | 33.4 |
| sin_Sinh | 33.4 |
| pan_Guru | 33.1 |
| npi_Deva | 32.9 |
| ckb_Arab | 32.8 |
| kaz_Cyrl | 32.4 |
| hau_Latn | 32.1 |
| hye_Armn | 31.9 |
| mya_Mymr | 31.3 |
| khk_Cyrl | 31.1 |
| guj_Gujr | 31.1 |
| lin_Latn | 31.0 |
| lug_Latn | 30.9 |
| ssw_Latn | 30.7 |
| khm_Khmr | 30.6 |
| plt_Latn | 30.5 |
| ben_Latn | 30.4 |
| som_Latn | 30.3 |
| pbt_Arab | 30.2 |
| zul_Latn | 30.2 |
| nso_Latn | 30.1 |
| tsn_Latn | 30.1 |
| yor_Latn | 30.1 |
| ibo_Latn | 30.1 |
| mal_Mlym | 30.1 |
| xho_Latn | 29.9 |
| fuv_Latn | 29.8 |
| gaz_Latn | 29.3 |
| ory_Orya | 29.2 |
| amh_Ethi | 28.9 |
| wol_Latn | 28.9 |
| tel_Telu | 27.5 |
| lao_Laoo | 26.5 |
| kan_Knda | 21.9 |

### Translation and instruction-language comparisons

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

The first two rows cover 91 non-English languages; the instruction-language comparison covers 89. These subsets are distinct from the full 122-language result.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Model | Variant | Evaluation setting | Average accuracy (%) | Language share ≥50% accuracy (%) | Language share ≥70% accuracy (%) | English accuracy (%) |
| --- | --- | --- | --- | --- | --- | --- |
| Llama-2-chat | 70B | In-Language | 44.1 | 35.2% | 2.2% | 78.8 |

## Other Belebele authors Models

- [BLOOMZ 7.1B (Belebele zero-shot)](/models/native-belebele-bloomz-7-1b-zero-shot) - Score: not computed
- [Falcon 40B (Belebele 5-shot)](/models/native-belebele-falcon-40b-5-shot) - Score: not computed
- [GPT3.5-turbo unk (Belebele zero-shot)](/models/native-belebele-gpt3-5-turbo-unk-zero-shot) - Score: not computed
- [InfoXLM large (550M) (Belebele English fine-tuning)](/models/native-belebele-infoxlm-large-550m-english-fine-tuning) - Score: not computed
- [InfoXLM large (550M) (Belebele Translate-Train-All)](/models/native-belebele-infoxlm-large-550m-translate-train-all) - Score: not computed
- [Llama 1 13B (Belebele 5-shot)](/models/native-belebele-llama-1-13b-5-shot) - Score: not computed
- [Llama 1 30B (Belebele 5-shot)](/models/native-belebele-llama-1-30b-5-shot) - Score: not computed
- [Llama 1 65B (Belebele 5-shot)](/models/native-belebele-llama-1-65b-5-shot) - Score: not computed
- [Llama 1 7B (Belebele 5-shot)](/models/native-belebele-llama-1-7b-5-shot) - Score: not computed
- [Llama 2 base 70B (Belebele 5-shot)](/models/native-belebele-llama-2-base-70b-5-shot) - Score: not computed
- [Llama-2-chat 70B (Belebele 89-language English instructions)](/models/native-belebele-llama-2-chat-70b-89-language-english-instructions) - Score: not computed
- [Llama-2-chat 70B (Belebele 89-language translated instructions)](/models/native-belebele-llama-2-chat-70b-89-language-translated-instructions) - Score: not computed
- [Llama-2-chat 70B (Belebele 91-language Translate-Test)](/models/native-belebele-llama-2-chat-70b-91-language-translate-test) - Score: not computed
- [Llama-2-chat 70B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-70b-zero-shot) - Score: not computed
- [Llama-2-chat 7B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-7b-zero-shot) - Score: not computed
- [XLM-R large (550M) (Belebele English fine-tuning)](/models/native-belebele-xlm-r-large-550m-english-fine-tuning) - Score: not computed
- [XLM-R large (550M) (Belebele Translate-Train-All)](/models/native-belebele-xlm-r-large-550m-translate-train-all) - Score: not computed
- [XLM-V large (1.2B) (Belebele English fine-tuning)](/models/native-belebele-xlm-v-large-1-2b-english-fine-tuning) - Score: not computed
- [XLM-V large (1.2B) (Belebele Translate-Train-All)](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) - Score: not computed
