# Belebele

> Belebele tests passage comprehension through questions with four answer options. Its parallel language variants let researchers study how reading performance changes across languages.

Canonical page: https://benchlm.ai/benchmarks/belebele

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026 source review

## About Belebele

- Year: 2023
- Tasks: Multilingual multiple-choice reading comprehension
- Format: Decision classification
- Difficulty: Depends on the evaluated split and label mapping
- Paper: [The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants](https://github.com/facebookresearch/belebele)

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

Belebele is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Original benchmark results

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

[Download full results (JSON)](/api/data/decision-benchmarks?benchmark=belebele)

### Paper summary across 122 language variants

Source vocabulary and parameter sizes are metadata, not scores. Source summary averages remain unchanged despite appendix discrepancies.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884)

| Model | Size or variant | Vocabulary size | Average accuracy (%) | Language share ≥50% accuracy (%) | Language share ≥70% accuracy (%) | English accuracy (%) | Non-English average (%) |
| --- | --- | --- | --- | --- | --- | --- | --- |
| 5-Shot In-Context Learning (examples in English) |  |  |  |  |  |  |  |
| [Llama 1](/models/native-belebele-llama-1-7b-5-shot) | 7B | 32K | 27.7 | 0.0% | 0.0% | 37.3 | 27.6 |
| [Llama 1](/models/native-belebele-llama-1-13b-5-shot) | 13B | 32K | 30.4 | 0.8% | 0.0% | 53.3 | 30.2 |
| [Llama 1](/models/native-belebele-llama-1-30b-5-shot) | 30B | 32K | 36.2 | 18.0% | 0.8% | 73.1 | 35.9 |
| [Llama 1](/models/native-belebele-llama-1-65b-5-shot) | 70B | 32K | 40.9 | 25.4% | 12.3% | 82.5 | 40.5 |
| [Llama 2 base](/models/native-belebele-llama-2-base-70b-5-shot) | 70B | 32K | 48.0 | 38.5% | 26.2% | 90.9 | 47.7 |
| [Falcon](/models/native-belebele-falcon-40b-5-shot) | 40B | 65K | 37.3 | 16.4% | 1.6% | 77.2 | 36.9 |
| Zero-Shot for Instructed Models (English instructions) |  |  |  |  |  |  |  |
| [BLOOMZ**](/models/native-belebele-bloomz-7-1b-zero-shot) | 7.1B | 251K | 43.2 | 28.7% | 9.0% | 79.6 | 42.9 |
| [Llama-2-chat](/models/native-belebele-llama-2-chat-7b-zero-shot) | 7B | 32K | 34.4 | 4.1% | 0.0% | 58.6 | 34.1 |
| [Llama-2-chat](/models/native-belebele-llama-2-chat-70b-zero-shot) | 70B | 32K | 41.5 | 27.0% | 2.5% | 78.8 | 41.2 |
| [GPT3.5-turbo](/models/native-belebele-gpt3-5-turbo-unk-zero-shot) | unk | 100K | 51.1 | 44.2% | 29.2% | 87.7 | 50.7 |
| Full Finetuning in English |  |  |  |  |  |  |  |
| [XLM-R](/models/native-belebele-xlm-r-large-550m-english-fine-tuning) | large (550M) | 250K | 54.0 | 64.8% | 15.6% | 76.2 | 53.8 |
| [XLM-V](/models/native-belebele-xlm-v-large-1-2b-english-fine-tuning) | large (1.2B) | 902K | 55.6 | 69.7% | 21.2% | 76.2 | 54.9 |
| [InfoXLM](/models/native-belebele-infoxlm-large-550m-english-fine-tuning) | large (550M) | 250K | 56.2 | 67.2% | 28.7% | 79.3 | 56.0 |
| Translate-Train-All |  |  |  |  |  |  |  |
| [XLM-R](/models/native-belebele-xlm-r-large-550m-translate-train-all) | large (550M) | 250K | 58.9 | 69.7% | 36.1% | 78.7 | 58.8 |
| [XLM-V](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) | large (1.2B) | 902K | 60.2 | 76.2% | 32.8% | 77.8 | 60.1 |
| [InfoXLM](/models/native-belebele-infoxlm-large-550m-translate-train-all) | large (550M) | 250K | 60.0 | 70.5% | 36.9% | 81.2 | 59.8 |

### All 122 language results: multilingual encoders

Accuracy values are percentages. PCT rows are the percentage of language variants meeting an accuracy threshold; summary and appendix PCT values differ and are both retained.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884)

| Language or summary metric | [XLM-V large (1.2B) (Belebele English fine-tuning)](/models/native-belebele-xlm-v-large-1-2b-english-fine-tuning) | [InfoXLM large (550M) (Belebele English fine-tuning)](/models/native-belebele-infoxlm-large-550m-english-fine-tuning) | [XLM-R large (550M) (Belebele English fine-tuning)](/models/native-belebele-xlm-r-large-550m-english-fine-tuning) | [XLM-V large (1.2B) (Belebele Translate-Train-All)](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) | [InfoXLM large (550M) (Belebele Translate-Train-All)](/models/native-belebele-infoxlm-large-550m-translate-train-all) | [XLM-R large (550M) (Belebele Translate-Train-All)](/models/native-belebele-xlm-r-large-550m-translate-train-all) |
| --- | --- | --- | --- | --- | --- | --- |
| AVG | 55.6 | 56.2 | 54.0 | 60.2 | 60.0 | 58.9 |
| PCT Above 50 | 69.7% | 67.2% | 64.8% | 76.2% | 70.5% | 69.7% |
| PCT Above 70 | 21.9% | 28.9% | 15.7% | 33.1% | 37.2% | 36.4% |
| eng_Latn | 76.2 | 79.3 | 76.2 | 77.8 | 81.2 | 78.7 |
| acm_Arab | 51.2 | 57.3 | 55.4 | 55.3 | 57.6 | 59.2 |
| afr_Latn | 69.3 | 72.7 | 69.1 | 72.3 | 75.1 | 74.3 |
| als_Latn | 68.4 | 68.9 | 64.9 | 70.8 | 72.2 | 71.4 |
| amh_Ethi | 53.1 | 52.9 | 52.6 | 61.6 | 60.0 | 60.7 |
| apc_Arab | 56.1 | 58.8 | 57.9 | 57.7 | 60.6 | 61.9 |
| arb_Arab | 67.2 | 71.0 | 69.8 | 70.6 | 75.0 | 74.3 |
| arb_Latn | 29.3 | 32.2 | 27.6 | 31.6 | 33.4 | 30.6 |
| ars_Arab | 55.6 | 59.9 | 58.9 | 61.1 | 65.8 | 65.9 |
| ary_Arab | 43.8 | 48.7 | 44.0 | 48.0 | 52.8 | 52.6 |
| arz_Arab | 56.9 | 60.2 | 57.6 | 61.4 | 64.9 | 66.1 |
| asm_Beng | 53.7 | 53.6 | 49.3 | 58.6 | 58.8 | 56.9 |
| azj_Latn | 59.7 | 61.3 | 59.0 | 65.0 | 65.6 | 65.1 |
| bam_Latn | 34.2 | 34.9 | 33.2 | 39.2 | 39.1 | 36.9 |
| ben_Beng | 60.0 | 63.4 | 59.6 | 65.6 | 69.6 | 63.7 |
| ben_Latn | 46.8 | 36.9 | 38.8 | 53.0 | 42.7 | 48.1 |
| bod_Tibt | 24.0 | 24.9 | 23.7 | 24.8 | 23.3 | 36.9 |
| bul_Cyrl | 72.6 | 72.0 | 70.1 | 74.0 | 75.3 | 74.2 |
| cat_Latn | 71.6 | 74.4 | 72.0 | 75.7 | 78.1 | 74.7 |
| ceb_Latn | 45.4 | 44.1 | 42.3 | 52.0 | 52.6 | 50.7 |
| ces_Latn | 69.9 | 72.3 | 69.9 | 72.3 | 76.2 | 74.4 |
| ckb_Arab | 29.7 | 52.3 | 30.3 | 36.9 | 58.0 | 36.9 |
| dan_Latn | 70.8 | 74.1 | 72.9 | 73.0 | 76.3 | 74.7 |
| deu_Latn | 72.6 | 75.7 | 72.9 | 74.1 | 78.7 | 76.7 |
| ell_Grek | 70.3 | 72.3 | 70.3 | 73.1 | 74.9 | 73.0 |
| est_Latn | 63.2 | 67.2 | 64.8 | 68.7 | 70.7 | 70.4 |
| eus_Latn | 63.6 | 66.1 | 64.8 | 68.2 | 70.8 | 70.3 |
| fin_Latn | 69.1 | 72.4 | 72.2 | 73.0 | 75.2 | 74.9 |
| fra_Latn | 73.1 | 74.2 | 72.1 | 74.6 | 76.8 | 75.6 |
| fuv_Latn | 29.7 | 27.7 | 26.4 | 32.8 | 30.7 | 31.1 |
| gaz_Latn | 48.8 | 33.8 | 36.4 | 52.6 | 36.0 | 43.3 |
| grn_Latn | 53.9 | 37.8 | 37.9 | 59.6 | 40.6 | 41.9 |
| guj_Gujr | 58.7 | 57.0 | 54.1 | 63.3 | 65.9 | 63.1 |
| hat_Latn | 57.1 | 39.6 | 35.2 | 63.2 | 44.1 | 39.8 |
| hau_Latn | 51.0 | 41.1 | 48.2 | 53.4 | 48.1 | 53.0 |
| heb_Hebr | 67.2 | 68.2 | 64.8 | 69.3 | 72.3 | 70.6 |
| hin_Deva | 57.9 | 60.2 | 57.4 | 63.8 | 64.3 | 63.4 |
| hin_Latn | 53.1 | 49.7 | 46.8 | 57.6 | 55.4 | 58.9 |
| hrv_Latn | 70.0 | 72.4 | 69.9 | 71.2 | 75.3 | 74.0 |
| hun_Latn | 69.7 | 70.8 | 70.0 | 73.1 | 74.2 | 72.8 |
| hye_Armn | 59.4 | 61.0 | 58.9 | 65.9 | 66.1 | 64.7 |
| ibo_Latn | 40.1 | 32.2 | 31.2 | 46.8 | 32.0 | 32.2 |
| ilo_Latn | 37.4 | 36.3 | 33.8 | 38.1 | 40.6 | 39.7 |
| ind_Latn | 68.9 | 70.7 | 68.0 | 71.3 | 73.1 | 70.4 |
| isl_Latn | 67.3 | 66.0 | 63.8 | 70.1 | 68.9 | 69.0 |
| ita_Latn | 70.6 | 72.8 | 70.0 | 71.8 | 76.4 | 73.3 |
| jav_Latn | 64.2 | 59.8 | 60.8 | 67.2 | 63.3 | 66.8 |
| jpn_Jpan | 66.4 | 70.1 | 67.6 | 71.3 | 71.8 | 71.0 |
| kac_Latn | 32.0 | 29.1 | 32.1 | 33.8 | 34.0 | 33.3 |
| kan_Knda | 61.1 | 62.0 | 59.7 | 66.6 | 68.4 | 69.1 |
| kat_Geor | 64.7 | 64.8 | 63.6 | 68.0 | 68.9 | 67.4 |
| kaz_Cyrl | 60.1 | 61.6 | 56.8 | 64.9 | 65.3 | 64.7 |
| kea_Latn | 44.0 | 45.2 | 44.9 | 48.7 | 47.7 | 48.1 |
| khk_Cyrl | 56.7 | 58.8 | 57.8 | 61.1 | 64.6 | 64.2 |
| khm_Khmr | 60.0 | 59.0 | 57.7 | 63.0 | 64.2 | 63.8 |
| kin_Latn | 35.9 | 33.6 | 34.3 | 39.1 | 39.1 | 38.6 |
| kir_Cyrl | 65.4 | 63.4 | 61.8 | 68.3 | 68.2 | 67.7 |
| kor_Hang | 70.1 | 71.4 | 68.7 | 72.9 | 74.6 | 74.8 |
| lao_Laoo | 55.8 | 57.6 | 53.0 | 63.2 | 63.6 | 63.0 |
| lin_Latn | 44.7 | 33.2 | 30.6 | 50.9 | 35.3 | 34.4 |
| lit_Latn | 68.3 | 69.4 | 67.2 | 71.7 | 72.9 | 72.0 |
| lug_Latn | 39.9 | 29.4 | 31.6 | 47.8 | 34.7 | 34.7 |
| luo_Latn | 30.3 | 30.9 | 30.8 | 33.7 | 34.9 | 33.2 |
| lvs_Latn | 70.1 | 71.3 | 68.7 | 74.1 | 75.6 | 73.0 |
| mal_Mlym | 62.0 | 65.0 | 62.7 | 69.1 | 68.3 | 67.1 |
| mar_Deva | 62.6 | 65.2 | 60.8 | 69.2 | 68.8 | 67.2 |
| mkd_Cyrl | 67.8 | 69.3 | 65.7 | 71.0 | 73.8 | 72.8 |
| mlt_Latn | 37.9 | 57.1 | 38.1 | 40.2 | 63.7 | 42.7 |
| mri_Latn | 32.0 | 30.6 | 32.2 | 33.0 | 35.7 | 34.0 |
| mya_Mymr | 56.6 | 59.1 | 53.6 | 62.2 | 65.1 | 62.9 |
| nld_Latn | 68.4 | 71.7 | 71.0 | 68.6 | 74.0 | 72.8 |
| nob_Latn | 71.8 | 73.6 | 70.7 | 72.8 | 75.4 | 74.2 |
| npi_Deva | 58.4 | 60.7 | 55.7 | 64.4 | 65.8 | 62.7 |
| npi_Latn | 38.3 | 35.8 | 33.8 | 37.4 | 36.4 | 34.8 |
| nso_Latn | 45.9 | 31.3 | 30.0 | 53.2 | 34.1 | 34.7 |
| nya_Latn | 31.0 | 29.2 | 29.8 | 34.2 | 33.0 | 30.8 |
| ory_Orya | 60.8 | 62.1 | 58.6 | 65.6 | 65.4 | 63.9 |
| pan_Guru | 58.1 | 59.2 | 57.8 | 63.1 | 62.6 | 62.0 |
| pbt_Arab | 55.4 | 56.0 | 51.0 | 60.6 | 62.6 | 61.1 |
| pes_Arab | 68.3 | 69.1 | 68.2 | 70.8 | 73.6 | 72.0 |
| plt_Latn | 55.7 | 45.6 | 52.7 | 61.7 | 53.4 | 58.1 |
| pol_Latn | 69.0 | 70.4 | 67.4 | 72.1 | 73.7 | 72.7 |
| por_Latn | 70.9 | 74.3 | 70.6 | 73.8 | 77.1 | 74.0 |
| ron_Latn | 72.3 | 72.9 | 71.3 | 74.0 | 76.2 | 74.8 |
| rus_Cyrl | 71.9 | 73.8 | 72.2 | 75.4 | 76.8 | 77.1 |
| shn_Mymr | 26.9 | 25.2 | 26.3 | 25.0 | 26.4 | 27.0 |
| sin_Latn | 24.9 | 34.2 | 30.7 | 41.7 | 38.3 | 37.3 |
| sin_Sinh | 64.4 | 67.2 | 62.7 | 69.8 | 70.2 | 68.6 |
| slk_Latn | 69.3 | 71.9 | 70.2 | 72.6 | 76.7 | 73.0 |
| slv_Latn | 69.7 | 72.2 | 68.6 | 71.8 | 75.4 | 73.9 |
| sna_Latn | 34.8 | 37.2 | 33.2 | 37.1 | 38.6 | 35.9 |
| snd_Arab | 55.2 | 56.6 | 51.9 | 60.0 | 61.3 | 61.3 |
| som_Latn | 46.0 | 39.1 | 42.6 | 50.7 | 46.3 | 50.7 |
| sot_Latn | 46.8 | 29.3 | 31.3 | 52.0 | 31.9 | 32.7 |
| spa_Latn | 71.0 | 73.3 | 71.4 | 72.7 | 75.3 | 76.4 |
| srp_Cyrl | 71.0 | 70.9 | 71.1 | 73.6 | 76.1 | 75.9 |
| ssw_Latn | 39.8 | 30.6 | 34.3 | 47.1 | 34.3 | 38.9 |
| sun_Latn | 60.9 | 50.7 | 55.3 | 64.2 | 55.8 | 59.4 |
| swe_Latn | 73.0 | 75.0 | 74.2 | 74.2 | 76.9 | 75.1 |
| swh_Latn | 64.9 | 65.3 | 62.8 | 69.3 | 69.2 | 68.7 |
| tam_Taml | 61.8 | 64.6 | 61.7 | 67.4 | 69.4 | 65.3 |
| tel_Telu | 55.6 | 57.8 | 53.6 | 62.1 | 63.2 | 61.1 |
| tgk_Cyrl | 38.2 | 58.6 | 33.8 | 39.2 | 64.3 | 39.6 |
| tgl_Latn | 69.2 | 67.4 | 64.7 | 72.0 | 70.4 | 70.0 |
| tha_Thai | 63.8 | 68.1 | 65.8 | 69.0 | 68.9 | 70.1 |
| tir_Ethi | 33.3 | 36.7 | 33.8 | 39.9 | 42.1 | 37.7 |
| tsn_Latn | 49.0 | 35.0 | 30.8 | 49.8 | 35.7 | 34.3 |
| tso_Latn | 37.9 | 36.3 | 34.2 | 41.7 | 39.7 | 37.1 |
| tur_Latn | 66.7 | 70.2 | 66.8 | 70.6 | 72.0 | 72.0 |
| ukr_Cyrl | 70.4 | 70.9 | 71.0 | 72.3 | 74.9 | 75.0 |
| urd_Arab | 61.6 | 63.8 | 59.3 | 65.6 | 68.6 | 66.3 |
| urd_Latn | 42.2 | 42.6 | 40.8 | 49.4 | 48.9 | 48.4 |
| uzn_Latn | 65.2 | 66.9 | 64.4 | 69.1 | 70.6 | 70.2 |
| vie_Latn | 69.6 | 71.1 | 69.4 | 73.7 | 72.9 | 71.4 |
| war_Latn | 46.4 | 44.7 | 43.7 | 47.6 | 49.3 | 46.6 |
| wol_Latn | 36.8 | 32.2 | 30.4 | 40.6 | 32.3 | 32.2 |
| xho_Latn | 48.7 | 36.1 | 39.0 | 54.4 | 40.2 | 45.4 |
| yor_Latn | 35.0 | 29.3 | 28.7 | 38.6 | 32.0 | 27.9 |
| zho_Hans | 69.8 | 74.6 | 71.0 | 73.7 | 76.2 | 74.8 |
| zho_Hant | 69.2 | 72.4 | 67.1 | 73.1 | 74.3 | 71.3 |
| zsm_Latn | 69.1 | 72.6 | 69.9 | 72.4 | 73.3 | 72.2 |
| zul_Latn | 46.9 | 36.4 | 39.0 | 54.2 | 39.8 | 44.1 |

### Multilingual encoder training settings

These are training settings, not benchmark scores. XLM-R Translate-Train-All learning rate is printed as 3-6 in the source; no exponent is inferred.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884)

| Training setting | [XLM-V large (1.2B) (Belebele English fine-tuning)](/models/native-belebele-xlm-v-large-1-2b-english-fine-tuning) | [InfoXLM large (550M) (Belebele English fine-tuning)](/models/native-belebele-infoxlm-large-550m-english-fine-tuning) | [XLM-R large (550M) (Belebele English fine-tuning)](/models/native-belebele-xlm-r-large-550m-english-fine-tuning) | [XLM-V large (1.2B) (Belebele Translate-Train-All)](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) | [InfoXLM large (550M) (Belebele Translate-Train-All)](/models/native-belebele-infoxlm-large-550m-translate-train-all) | [XLM-R large (550M) (Belebele Translate-Train-All)](/models/native-belebele-xlm-r-large-550m-translate-train-all) |
| --- | --- | --- | --- | --- | --- | --- |
| epochs | 3 | 4 | 3 | 1 | 1 | 1 |
| training set size | 67.5k | 67.5k | 67.5k | 650k | 650k | 650k |
| learning rate | 5e-6 | 4e-6 | 5e-6 | 3-6 | 3e-6 | 3e-6 |
| weight decay | 0.01 | 0.01 | 0.01 | 0.001 | 0.001 | 0.001 |
| batch size | 64 | 64 | 64 | 64 | 64 | 64 |

### All 122 language results: large language models

The GPT-3.5 appendix average is 50.6; the summary prints 51.1. The Llama 1 checkpoint is labeled 65B in this appendix.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884)

| Language or summary metric | [GPT3.5-turbo unk (Belebele zero-shot)](/models/native-belebele-gpt3-5-turbo-unk-zero-shot) | [Llama-2-chat 70B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-70b-zero-shot) | [Llama 2 base 70B (Belebele 5-shot)](/models/native-belebele-llama-2-base-70b-5-shot) | [Llama 1 65B (Belebele 5-shot)](/models/native-belebele-llama-1-65b-5-shot) | [Falcon 40B (Belebele 5-shot)](/models/native-belebele-falcon-40b-5-shot) | [XLM-V large (1.2B) (Belebele Translate-Train-All)](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) |
| --- | --- | --- | --- | --- | --- | --- |
| AVG | 50.6 | 41.5 | 48.0 | 40.9 | 37.3 | 60.2 |
| PCT Above 50 | 43.4 % | 27.1% | 38.5 % | 25.4% | 16.4% | 76.2% |
| PCT Above 70 | 28.9 % | 2.5% | 26.2 % | 12.3% | 1.6% | 33.1% |
| eng_Latn | 87.7 | 78.8 | 90.9 | 82.5 | 77.2 | 77.8 |
| acm_Arab | 51.6 | 35.9 | 47.9 | 37.9 | 37.6 | 55.3 |
| afr_Latn | 78.3 | 57.9 | 75.9 | 60.7 | 53.4 | 72.3 |
| als_Latn | 67.1 | 36.0 | 45.4 | 34.9 | 36.6 | 70.8 |
| amh_Ethi | 28.7 | 28.9 | 27.5 | 27.8 | 24.8 | 61.6 |
| apc_Arab | 55.6 | 38.8 | 51.2 | 39.6 | 36.3 | 57.7 |
| arb_Arab | 69.3 | 42.3 | 61.7 | 44.1 | 38.3 | 70.6 |
| arb_Latn | 31.1 | 30.2 | 26.8 | 28.0 | 26.3 | 31.6 |
| ars_Arab | 55.1 | 37.4 | 50.2 | 40.7 | 32.1 | 61.1 |
| ary_Arab | 45.7 | 32.6 | 40.6 | 33.1 | 32.3 | 48.0 |
| arz_Arab | 56.7 | 37.3 | 50.7 | 37.4 | 33.0 | 61.4 |
| asm_Beng | 36.0 | 35.7 | 32.3 | 28.9 | 22.4 | 58.6 |
| azj_Latn | 54.9 | 33.4 | 42.2 | 33.6 | 34.1 | 65.0 |
| bam_Latn | 31.7 | 29.4 | 30.3 | 28.4 | 29.7 | 39.2 |
| ben_Beng | 43.6 | 34.9 | 39.1 | 33.4 | 22.6 | 65.6 |
| ben_Latn | 34.6 | 30.4 | 29.6 | 29.2 | 32.1 | 53.0 |
| bod_Tibt | 26.6 | 28.3 | 25.7 | 24.9 | 26.8 | 24.8 |
| bul_Cyrl | 76.0 | 65.0 | 80.4 | 69.3 | 41.9 | 74.0 |
| cat_Latn | 78.4 | 68.2 | 84.6 | 76.3 | 58.8 | 75.7 |
| ceb_Latn | 53.3 | 40.6 | 50.4 | 38.9 | 39.2 | 52.0 |
| ces_Latn | 76.9 | 65.0 | 81.1 | 70.7 | 65.0 | 72.3 |
| ckb_Arab | 31.8 | 32.8 | 28.7 | 31.6 | 28.9 | 36.9 |
| dan_Latn | 80.7 | 66.2 | 83.6 | 73.6 | 56.2 | 73.0 |
| deu_Latn | 83.3 | 69.4 | 84.6 | 76.0 | 70.1 | 74.1 |
| ell_Grek | 73.0 | 50.7 | 64.9 | 44.2 | 31.2 | 73.1 |
| est_Latn | 73.1 | 36.6 | 53.0 | 36.3 | 34.9 | 68.7 |
| eus_Latn | 40.9 | 31.1 | 34.7 | 32.8 | 38.9 | 68.2 |
| fin_Latn | 77.9 | 62.7 | 79.3 | 55.7 | 42.8 | 73.0 |
| fra_Latn | 83.1 | 72.2 | 86.4 | 77.5 | 69.7 | 74.6 |
| fuv_Latn | 26.1 | 29.8 | 24.9 | 25.4 | 25.1 | 32.8 |
| gaz_Latn | 30.3 | 29.3 | 27.8 | 29.1 | 24.9 | 52.6 |
| grn_Latn | 34.2 | 32.2 | 32.4 | 30.3 | 33.8 | 59.6 |
| guj_Gujr | 38.4 | 31.1 | 27.1 | 25.7 | 24.7 | 63.3 |
| hat_Latn | 51.6 | 34.1 | 37.4 | 33.7 | 36.2 | 63.2 |
| hau_Latn | 32.2 | 32.1 | 28.0 | 26.4 | 28.9 | 53.4 |
| heb_Hebr | 64.2 | 41.4 | 54.9 | 41.4 | 31.1 | 69.3 |
| hin_Deva | 49.1 | 42.0 | 52.6 | 38.4 | 27.1 | 63.8 |
| hin_Latn | 52.3 | 39.2 | 49.0 | 34.2 | 40.0 | 57.6 |
| hrv_Latn | 78.4 | 64.7 | 79.8 | 66.9 | 48.7 | 71.2 |
| hun_Latn | 74.6 | 61.1 | 78.8 | 66.7 | 37.7 | 73.1 |
| hye_Armn | 35.0 | 31.9 | 34.1 | 32.1 | 25.4 | 65.9 |
| ibo_Latn | 28.4 | 30.1 | 27.4 | 25.3 | 30.2 | 46.8 |
| ilo_Latn | 37.1 | 33.2 | 36.6 | 32.1 | 35.1 | 38.1 |
| ind_Latn | 74.2 | 61.3 | 81.4 | 55.7 | 52.1 | 71.3 |
| isl_Latn | 62.3 | 38.0 | 54.3 | 42.1 | 36.4 | 70.1 |
| ita_Latn | 80.0 | 68.6 | 84.5 | 76.1 | 66.4 | 71.8 |
| jav_Latn | 46.7 | 37.0 | 40.3 | 33.0 | 36.8 | 67.2 |
| jpn_Jpan | 70.9 | 56.6 | 77.6 | 53.9 | 49.6 | 71.3 |
| kac_Latn | 30.9 | 30.7 | 27.7 | 28.6 | 27.8 | 33.8 |
| kan_Knda | 40.6 | 21.9 | 25.7 | 24.4 | 24.0 | 66.6 |
| kat_Geor | 33.0 | 34.6 | 37.8 | 34.3 | 23.4 | 68.0 |
| kaz_Cyrl | 35.0 | 32.4 | 29.3 | 32.4 | 32.6 | 64.9 |
| kea_Latn | 46.0 | 38.1 | 45.4 | 38.1 | 38.0 | 48.7 |
| khk_Cyrl | 32.0 | 31.1 | 29.8 | 28.4 | 27.4 | 61.1 |
| khm_Khmr | 30.4 | 30.6 | 27.0 | 28.2 | 25.0 | 63.0 |
| kin_Latn | 35.2 | 30.6 | 29.8 | 28.5 | 31.9 | 39.1 |
| kir_Cyrl | 37.9 | 32.2 | 34.6 | 32.5 | 31.9 | 68.3 |
| kor_Hang | 67.1 | 56.3 | 77.8 | 52.9 | 40.2 | 72.9 |
| lao_Laoo | 30.0 | 26.5 | 24.3 | 26.2 | 28.1 | 63.2 |
| lin_Latn | 33.8 | 31.0 | 28.0 | 30.4 | 29.3 | 50.9 |
| lit_Latn | 72.0 | 39.7 | 52.1 | 39.6 | 39.3 | 71.7 |
| lug_Latn | 28.4 | 30.9 | 29.2 | 28.3 | 28.9 | 47.8 |
| luo_Latn | 27.1 | 31.2 | 29.4 | 29.3 | 29.9 | 33.7 |
| lvs_Latn | 70.8 | 41.0 | 51.3 | 39.0 | 37.6 | 74.1 |
| mal_Mlym | 34.9 | 30.1 | 32.4 | 30.0 | 21.2 | 69.1 |
| mar_Deva | 38.3 | 34.8 | 41.2 | 32.9 | 25.0 | 69.2 |
| mkd_Cyrl | 69.4 | 55.7 | 72.5 | 56.2 | 38.1 | 71.0 |
| mlt_Latn | 44.8 | 36.2 | 44.9 | 36.7 | 35.4 | 40.2 |
| mri_Latn | 33.3 | 31.8 | 28.5 | 32.0 | 29.7 | 33.0 |
| mya_Mymr | 30.3 | 31.3 | 24.1 | 24.2 | 22.6 | 62.2 |
| nld_Latn | 80.4 | 66.2 | 82.2 | 73.3 | 66.7 | 68.6 |
| nob_Latn | 79.0 | 65.7 | 81.8 | 70.9 | 60.8 | 72.8 |
| npi_Deva | 40.4 | 32.9 | 40.4 | 33.0 | 25.4 | 64.4 |
| npi_Latn | 35.1 | 30.4 | 30.2 | 30.0 | 30.9 | 37.4 |
| nso_Latn | 33.6 | 30.1 | 30.4 | 27.4 | 29.3 | 53.2 |
| nya_Latn | 33.2 | 29.3 | 27.3 | 28.7 | 29.3 | 34.2 |
| ory_Orya |  | 29.2 | 24.8 | 23.9 | 23.7 | 65.6 |
| pan_Guru | 39.1 | 33.1 | 26.3 | 27.1 | 23.4 | 63.1 |
| pbt_Arab | 32.3 | 30.2 | 30.8 | 29.4 | 29.4 | 60.6 |
| pes_Arab | 61.8 | 41.8 | 53.9 | 41.0 | 35.9 | 70.8 |
| plt_Latn | 32.3 | 30.5 | 29.6 | 31.0 | 31.4 | 61.7 |
| pol_Latn | 74.7 | 61.7 | 79.2 | 67.0 | 59.9 | 72.1 |
| por_Latn | 83.0 | 70.2 | 86.1 | 75.4 | 68.3 | 73.8 |
| ron_Latn | 77.4 | 65.6 | 83.4 | 73.2 | 66.6 | 74.0 |
| rus_Cyrl | 78.4 | 67.0 | 82.7 | 73.1 | 48.1 | 75.4 |
| shn_Mymr |  | 28.2 | 25.6 | 22.7 | 24.0 | 25.0 |
| sin_Latn | 30.4 | 31.9 | 33.8 | 27.9 | 32.6 | 41.7 |
| sin_Sinh | 32.6 | 33.4 | 25.2 | 29.4 | 27.7 | 69.8 |
| slk_Latn | 77.3 | 58.8 | 75.2 | 60.4 | 57.0 | 72.6 |
| slv_Latn | 77.4 | 62.4 | 76.7 | 65.6 | 43.7 | 71.8 |
| sna_Latn | 35.4 | 30.2 | 27.4 | 28.3 | 31.6 | 37.1 |
| snd_Arab | 34.1 | 29.7 | 30.9 | 28.9 | 30.2 | 60.0 |
| som_Latn | 32.4 | 30.3 | 27.8 | 27.6 | 29.9 | 50.7 |
| sot_Latn | 33.9 | 30.0 | 28.9 | 26.8 | 29.9 | 52.0 |
| spa_Latn | 79.2 | 68.4 | 85.0 | 74.8 | 69.2 | 72.7 |
| srp_Cyrl | 74.8 | 65.1 | 81.0 | 70.7 | 40.2 | 73.6 |
| ssw_Latn | 32.0 | 30.7 | 27.7 | 28.0 | 30.1 | 47.1 |
| sun_Latn | 38.9 | 34.9 | 37.8 | 30.7 | 34.1 | 64.2 |
| swe_Latn | 81.7 | 67.4 | 82.7 | 73.7 | 67.3 | 74.2 |
| swh_Latn | 70.3 | 35.1 | 39.6 | 34.4 | 36.7 | 69.3 |
| tam_Taml | 32.8 | 34.4 | 33.2 | 31.6 | 24.4 | 67.4 |
| tel_Telu | 34.6 | 27.5 | 25.9 | 26.6 | 22.4 | 62.1 |
| tgk_Cyrl | 37.7 | 32.5 | 34.0 | 33.1 | 32.7 | 39.2 |
| tgl_Latn | 66.7 | 49.6 | 68.1 | 48.3 | 47.7 | 72.0 |
| tha_Thai | 55.7 | 38.9 | 46.2 | 35.0 | 33.0 | 69.0 |
| tir_Ethi | 28.4 | 29.6 | 24.5 | 23.5 | 25.0 | 39.9 |
| tsn_Latn | 31.8 | 30.1 | 28.5 | 24.7 | 31.2 | 49.8 |
| tso_Latn | 33.4 | 30.0 | 30.4 | 28.0 | 29.7 | 41.7 |
| tur_Latn | 69.9 | 47.3 | 65.4 | 42.1 | 39.6 | 70.6 |
| ukr_Cyrl | 72.8 | 65.7 | 80.8 | 69.7 | 41.9 | 72.3 |
| urd_Arab | 48.3 | 37.0 | 43.2 | 34.7 | 31.7 | 65.6 |
| urd_Latn | 40.3 | 34.1 | 38.0 | 30.1 | 34.2 | 49.4 |
| uzn_Latn | 44.1 | 33.1 | 35.1 | 30.6 | 33.1 | 69.1 |
| vie_Latn | 72.9 | 59.6 | 78.4 | 43.5 | 41.4 | 73.7 |
| war_Latn | 48.9 | 39.3 | 44.4 | 37.4 | 38.6 | 47.6 |
| wol_Latn | 29.0 | 28.9 | 27.6 | 26.0 | 26.8 | 40.6 |
| xho_Latn | 30.0 | 29.9 | 28.2 | 27.6 | 30.2 | 54.4 |
| yor_Latn | 29.1 | 30.1 | 28.3 | 27.7 | 27.2 | 38.6 |
| zho_Hans | 77.6 | 62.4 | 83.7 | 64.6 | 66.0 | 73.7 |
| zho_Hant | 76.3 | 59.3 | 82.0 | 57.7 | 62.2 | 73.1 |
| zsm_Latn | 74.0 | 56.4 | 76.3 | 51.7 | 51.3 | 72.4 |
| zul_Latn | 30.4 | 30.2 | 29.7 | 27.1 | 30.7 | 54.2 |

### Multiple-script language comparison

Per-language accuracy (%). The final AVG column is a cross-system summary, not an additional model.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884)

| Language | [GPT3.5-turbo unk (Belebele zero-shot)](/models/native-belebele-gpt3-5-turbo-unk-zero-shot) | [Llama 2 base 70B (Belebele 5-shot)](/models/native-belebele-llama-2-base-70b-5-shot) | [Falcon 40B (Belebele 5-shot)](/models/native-belebele-falcon-40b-5-shot) | [InfoXLM large (550M) (Belebele English fine-tuning)](/models/native-belebele-infoxlm-large-550m-english-fine-tuning) | Average across shown systems (%) |
| --- | --- | --- | --- | --- | --- |
| arb_Arab | 69.3 | 61.7 | 38.3 | 71.0 | 60.1 |
| arb_Latn | 31.1 | 26.8 | 26.3 | 32.2 | 29.1 |
| ben_Beng | 43.6 | 39.1 | 22.6 | 63.4 | 42.2 |
| ben_Latn | 34.6 | 29.6 | 32.1 | 36.9 | 33.3 |
| hin_Deva | 49.1 | 52.6 | 27.1 | 60.2 | 47.3 |
| hin_Latn | 52.3 | 49.0 | 40.0 | 49.7 | 47.8 |
| npi_Deva | 40.4 | 40.4 | 25.4 | 60.7 | 41.7 |
| npi_Latn | 35.1 | 30.2 | 30.9 | 35.8 | 33.0 |
| sin_Sinh | 32.6 | 25.2 | 27.7 | 67.2 | 38.2 |
| sin_Latn | 30.4 | 33.8 | 32.6 | 34.2 | 32.8 |
| urd_Arab | 48.3 | 43.2 | 31.7 | 63.8 | 46.7 |
| urd_Latn | 40.3 | 38.0 | 34.2 | 42.6 | 38.8 |
| zho_Hant | 76.3 | 82.0 | 62.2 | 72.4 | 73.3 |
| zho_Hans | 77.6 | 83.7 | 66.0 | 74.6 | 75.5 |

### All translated-test language results

The 91 non-English-language subset also shows an English reference row. Its averages are not the 122-language averages. In-language appendix average is 44.0; the summary prints 44.1.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884)

| Language or summary metric | [Llama-2-chat 70B (Belebele 91-language in-language)](/models/native-belebele-llama-2-chat-70b-91-language-in-language) | [Llama-2-chat 70B (Belebele 91-language Translate-Test)](/models/native-belebele-llama-2-chat-70b-91-language-translate-test) | [XLM-V large (1.2B) (Belebele Translate-Train-All)](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) |
| --- | --- | --- | --- |
| AVG | 44.0 | 57.1 | 64.9 |
| PCT Above 50 | 35.2% | 78.0% | 90.1% |
| PCT Above 70 | 2.2% | 2.2% | 42.9% |
| eng_Latn | 78.8 | 78.8 | 77.8 |
| fra_Latn | 72.2 | 70.6 | 73.1 |
| por_Latn | 70.2 | 69.9 | 70.9 |
| deu_Latn | 69.4 | 65.7 | 72.6 |
| ita_Latn | 68.6 | 66.1 | 70.6 |
| spa_Latn | 68.4 | 69.3 | 71.0 |
| cat_Latn | 68.2 | 67.0 | 71.6 |
| swe_Latn | 67.4 | 66.1 | 73.0 |
| rus_Cyrl | 67.0 | 67.3 | 71.9 |
| dan_Latn | 66.2 | 66.8 | 70.8 |
| nld_Latn | 66.2 | 67.2 | 68.4 |
| nob_Latn | 65.7 | 68.3 | 71.8 |
| ukr_Cyrl | 65.7 | 66.0 | 70.4 |
| ron_Latn | 65.6 | 67.0 | 72.3 |
| srp_Cyrl | 65.1 | 66.2 | 71.0 |
| bul_Cyrl | 65.0 | 67.7 | 72.6 |
| ces_Latn | 65.0 | 65.6 | 69.9 |
| hrv_Latn | 64.7 | 65.3 | 70.0 |
| fin_Latn | 62.7 | 61.1 | 69.1 |
| slv_Latn | 62.4 | 61.2 | 69.7 |
| zho_Hans | 62.4 | 71.2 | 69.8 |
| pol_Latn | 61.7 | 63.0 | 69.0 |
| ind_Latn | 61.3 | 64.8 | 68.9 |
| hun_Latn | 61.1 | 62.9 | 69.7 |
| vie_Latn | 59.6 | 59.4 | 69.6 |
| zho_Hant | 59.3 | 65.8 | 69.2 |
| slk_Latn | 58.8 | 66.2 | 69.3 |
| afr_Latn | 57.9 | 65.0 | 69.3 |
| jpn_Jpan | 56.6 | 54.8 | 66.4 |
| zsm_Latn | 56.4 | 67.0 | 69.1 |
| kor_Hang | 56.3 | 56.7 | 70.1 |
| mkd_Cyrl | 55.7 | 66.7 | 67.8 |
| ell_Grek | 50.7 | 67.6 | 70.3 |
| tgl_Latn | 49.6 | 62.2 | 69.2 |
| tur_Latn | 47.3 | 62.6 | 66.7 |
| arb_Arab | 42.3 | 60.7 | 67.2 |
| hin_Deva | 42.0 | 62.8 | 57.9 |
| pes_Arab | 41.8 | 59.6 | 68.3 |
| heb_Hebr | 41.4 | 62.0 | 67.2 |
| lvs_Latn | 41.0 | 60.9 | 70.1 |
| ceb_Latn | 40.6 | 62.6 | 45.4 |
| lit_Latn | 39.7 | 60.8 | 68.3 |
| hin_Latn | 39.2 | 52.7 | 53.1 |
| tha_Thai | 38.9 | 54.1 | 63.8 |
| isl_Latn | 38.0 | 58.1 | 67.3 |
| jav_Latn | 37.0 | 55.3 | 64.2 |
| urd_Arab | 37.0 | 59.4 | 61.6 |
| est_Latn | 36.6 | 59.4 | 63.2 |
| als_Latn | 36.0 | 63.1 | 68.4 |
| asm_Beng | 35.7 | 57.7 | 53.7 |
| swh_Latn | 35.1 | 57.8 | 64.9 |
| ben_Beng | 34.9 | 61.0 | 60.0 |
| sun_Latn | 34.9 | 50.8 | 60.9 |
| mar_Deva | 34.8 | 60.0 | 62.6 |
| kat_Geor | 34.6 | 57.7 | 64.7 |
| tam_Taml | 34.4 | 55.9 | 61.8 |
| urd_Latn | 34.1 | 43.0 | 42.2 |
| hat_Latn | 34.1 | 56.3 | 57.1 |
| azj_Latn | 33.4 | 55.6 | 59.7 |
| sin_Sinh | 33.4 | 57.7 | 64.4 |
| pan_Guru | 33.1 | 57.6 | 58.1 |
| npi_Deva | 32.9 | 62.0 | 58.4 |
| ckb_Arab | 32.8 | 51.3 | 29.7 |
| kaz_Cyrl | 32.4 | 53.2 | 60.1 |
| hau_Latn | 32.1 | 43.4 | 51.0 |
| hye_Armn | 31.9 | 58.0 | 59.4 |
| mya_Mymr | 31.3 | 46.6 | 56.6 |
| khk_Cyrl | 31.1 | 52.2 | 56.7 |
| guj_Gujr | 31.1 | 59.6 | 58.7 |
| lin_Latn | 31.0 | 40.3 | 44.7 |
| lug_Latn | 30.9 | 38.7 | 39.9 |
| ssw_Latn | 30.7 | 43.2 | 39.8 |
| khm_Khmr | 30.6 | 52.8 | 60.0 |
| plt_Latn | 30.5 | 46.7 | 55.7 |
| ben_Latn | 30.4 | 45.1 | 46.8 |
| som_Latn | 30.3 | 40.8 | 46.0 |
| pbt_Arab | 30.2 | 48.8 | 55.4 |
| zul_Latn | 30.2 | 44.4 | 46.9 |
| nso_Latn | 30.1 | 43.4 | 45.9 |
| tsn_Latn | 30.1 | 40.4 | 49.0 |
| yor_Latn | 30.1 | 37.7 | 35.0 |
| ibo_Latn | 30.1 | 35.3 | 40.1 |
| mal_Mlym | 30.1 | 63.0 | 62.0 |
| xho_Latn | 29.9 | 49.2 | 48.7 |
| fuv_Latn | 29.8 | 29.4 | 29.7 |
| gaz_Latn | 29.3 | 37.0 | 48.8 |
| ory_Orya | 29.2 | 57.8 | 60.8 |
| amh_Ethi | 28.9 | 50.4 | 53.1 |
| wol_Latn | 28.9 | 39.0 | 36.8 |
| tel_Telu | 27.5 | 54.3 | 55.6 |
| lao_Laoo | 26.5 | 47.4 | 55.8 |
| kan_Knda | 21.9 | 62.0 | 61.1 |

### Translation and instruction-language comparisons

The first two rows cover 91 non-English languages; the instruction-language comparison covers 89. These subsets are distinct from the full 122-language result.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884)

| Model | Variant | Evaluation setting | Average accuracy (%) | Language share ≥50% accuracy (%) | Language share ≥70% accuracy (%) | English accuracy (%) |
| --- | --- | --- | --- | --- | --- | --- |
| Translate-Test (English) on 91 non-English languages in Zero-Shot |  |  |  |  |  |  |
| [Llama-2-chat](/models/native-belebele-llama-2-chat-70b-91-language-translate-test) | 70B | Translate-Test | 57.1 | 78.0% | 2.2% | 78.8 |
| [Llama-2-chat](/models/native-belebele-llama-2-chat-70b-91-language-in-language) | 70B | In-Language | 44.1 | 35.2% | 2.2% | 78.8 |
| Translated Instructions in 89 non-English languages Zero-Shot |  |  |  |  |  |  |
| [Llama-2-chat](/models/native-belebele-llama-2-chat-70b-89-language-translated-instructions) | 70B | In-Language Translated Instructions | 38.7 | 36.0% | 7.9% | 78.8 |
| [Llama-2-chat](/models/native-belebele-llama-2-chat-70b-89-language-english-instructions) | 70B | English Instructions | 44.9 | 37.1% | 3.4% | 78.8 |

## FAQ

### Which Belebele results are included?

This page includes 7 numeric result and configuration tables from Belebele’s original paper and available benchmark-owner updates, covering 20 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

### Can I compare these Belebele scores with the Perplexity panel?

Compare Belebele scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

### Where can I download the Belebele results?

The results download on this page provides every imported Belebele table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
