# InfoXLM large (550M) (Belebele English fine-tuning) Benchmark Scores & Performance

> InfoXLM large (550M) (Belebele English fine-tuning) has published results in 4 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-belebele-infoxlm-large-550m-english-fine-tuning

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | Belebele authors |
| Source Type | Research system |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://arxiv.org/abs/2308.16884) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: InfoXLM large (550M) (Belebele English fine-tuning)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-belebele-infoxlm-large-550m-english-fine-tuning) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-belebele-infoxlm-large-550m-english-fine-tuning&format=csv)

### Paper summary across 122 language variants

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

Source vocabulary and parameter sizes are metadata, not scores. Source summary averages remain unchanged despite appendix discrepancies.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Model | Size or variant | Vocabulary size | Average accuracy (%) | Language share ≥50% accuracy (%) | Language share ≥70% accuracy (%) | English accuracy (%) | Non-English average (%) |
| --- | --- | --- | --- | --- | --- | --- | --- |
| InfoXLM | large (550M) | 250K | 56.2 | 67.2% | 28.7% | 79.3 | 56.0 |

### All 122 language results: multilingual encoders

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

Accuracy values are percentages. PCT rows are the percentage of language variants meeting an accuracy threshold; summary and appendix PCT values differ and are both retained.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Language or summary metric | InfoXLM large (550M) (Belebele English fine-tuning) |
| --- | --- |
| AVG | 56.2 |
| PCT Above 50 | 67.2% |
| PCT Above 70 | 28.9% |
| eng_Latn | 79.3 |
| acm_Arab | 57.3 |
| afr_Latn | 72.7 |
| als_Latn | 68.9 |
| amh_Ethi | 52.9 |
| apc_Arab | 58.8 |
| arb_Arab | 71.0 |
| arb_Latn | 32.2 |
| ars_Arab | 59.9 |
| ary_Arab | 48.7 |
| arz_Arab | 60.2 |
| asm_Beng | 53.6 |
| azj_Latn | 61.3 |
| bam_Latn | 34.9 |
| ben_Beng | 63.4 |
| ben_Latn | 36.9 |
| bod_Tibt | 24.9 |
| bul_Cyrl | 72.0 |
| cat_Latn | 74.4 |
| ceb_Latn | 44.1 |
| ces_Latn | 72.3 |
| ckb_Arab | 52.3 |
| dan_Latn | 74.1 |
| deu_Latn | 75.7 |
| ell_Grek | 72.3 |
| est_Latn | 67.2 |
| eus_Latn | 66.1 |
| fin_Latn | 72.4 |
| fra_Latn | 74.2 |
| fuv_Latn | 27.7 |
| gaz_Latn | 33.8 |
| grn_Latn | 37.8 |
| guj_Gujr | 57.0 |
| hat_Latn | 39.6 |
| hau_Latn | 41.1 |
| heb_Hebr | 68.2 |
| hin_Deva | 60.2 |
| hin_Latn | 49.7 |
| hrv_Latn | 72.4 |
| hun_Latn | 70.8 |
| hye_Armn | 61.0 |
| ibo_Latn | 32.2 |
| ilo_Latn | 36.3 |
| ind_Latn | 70.7 |
| isl_Latn | 66.0 |
| ita_Latn | 72.8 |
| jav_Latn | 59.8 |
| jpn_Jpan | 70.1 |
| kac_Latn | 29.1 |
| kan_Knda | 62.0 |
| kat_Geor | 64.8 |
| kaz_Cyrl | 61.6 |
| kea_Latn | 45.2 |
| khk_Cyrl | 58.8 |
| khm_Khmr | 59.0 |
| kin_Latn | 33.6 |
| kir_Cyrl | 63.4 |
| kor_Hang | 71.4 |
| lao_Laoo | 57.6 |
| lin_Latn | 33.2 |
| lit_Latn | 69.4 |
| lug_Latn | 29.4 |
| luo_Latn | 30.9 |
| lvs_Latn | 71.3 |
| mal_Mlym | 65.0 |
| mar_Deva | 65.2 |
| mkd_Cyrl | 69.3 |
| mlt_Latn | 57.1 |
| mri_Latn | 30.6 |
| mya_Mymr | 59.1 |
| nld_Latn | 71.7 |
| nob_Latn | 73.6 |
| npi_Deva | 60.7 |
| npi_Latn | 35.8 |
| nso_Latn | 31.3 |
| nya_Latn | 29.2 |
| ory_Orya | 62.1 |
| pan_Guru | 59.2 |
| pbt_Arab | 56.0 |
| pes_Arab | 69.1 |
| plt_Latn | 45.6 |
| pol_Latn | 70.4 |
| por_Latn | 74.3 |
| ron_Latn | 72.9 |
| rus_Cyrl | 73.8 |
| shn_Mymr | 25.2 |
| sin_Latn | 34.2 |
| sin_Sinh | 67.2 |
| slk_Latn | 71.9 |
| slv_Latn | 72.2 |
| sna_Latn | 37.2 |
| snd_Arab | 56.6 |
| som_Latn | 39.1 |
| sot_Latn | 29.3 |
| spa_Latn | 73.3 |
| srp_Cyrl | 70.9 |
| ssw_Latn | 30.6 |
| sun_Latn | 50.7 |
| swe_Latn | 75.0 |
| swh_Latn | 65.3 |
| tam_Taml | 64.6 |
| tel_Telu | 57.8 |
| tgk_Cyrl | 58.6 |
| tgl_Latn | 67.4 |
| tha_Thai | 68.1 |
| tir_Ethi | 36.7 |
| tsn_Latn | 35.0 |
| tso_Latn | 36.3 |
| tur_Latn | 70.2 |
| ukr_Cyrl | 70.9 |
| urd_Arab | 63.8 |
| urd_Latn | 42.6 |
| uzn_Latn | 66.9 |
| vie_Latn | 71.1 |
| war_Latn | 44.7 |
| wol_Latn | 32.2 |
| xho_Latn | 36.1 |
| yor_Latn | 29.3 |
| zho_Hans | 74.6 |
| zho_Hant | 72.4 |
| zsm_Latn | 72.6 |
| zul_Latn | 36.4 |

### Multilingual encoder training settings

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

These are training settings, not benchmark scores. XLM-R Translate-Train-All learning rate is printed as 3-6 in the source; no exponent is inferred.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Training setting | InfoXLM large (550M) (Belebele English fine-tuning) |
| --- | --- |
| epochs | 4 |
| training set size | 67.5k |
| learning rate | 4e-6 |
| weight decay | 0.01 |
| batch size | 64 |

### Multiple-script language comparison

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

Per-language accuracy (%). The final AVG column is a cross-system summary, not an additional model.

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. [CC-BY-SA-4.0](https://creativecommons.org/licenses/by-sa/4.0/). Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2308.16884) · [Full Belebele results](/benchmarks/belebele)

| Language | InfoXLM large (550M) (Belebele English fine-tuning) |
| --- | --- |
| arb_Arab | 71.0 |
| arb_Latn | 32.2 |
| ben_Beng | 63.4 |
| ben_Latn | 36.9 |
| hin_Deva | 60.2 |
| hin_Latn | 49.7 |
| npi_Deva | 60.7 |
| npi_Latn | 35.8 |
| sin_Sinh | 67.2 |
| sin_Latn | 34.2 |
| urd_Arab | 63.8 |
| urd_Latn | 42.6 |
| zho_Hant | 72.4 |
| zho_Hans | 74.6 |

## Other Belebele authors Models

- [BLOOMZ 7.1B (Belebele zero-shot)](/models/native-belebele-bloomz-7-1b-zero-shot) - Score: not computed
- [Falcon 40B (Belebele 5-shot)](/models/native-belebele-falcon-40b-5-shot) - Score: not computed
- [GPT3.5-turbo unk (Belebele zero-shot)](/models/native-belebele-gpt3-5-turbo-unk-zero-shot) - Score: not computed
- [InfoXLM large (550M) (Belebele Translate-Train-All)](/models/native-belebele-infoxlm-large-550m-translate-train-all) - Score: not computed
- [Llama 1 13B (Belebele 5-shot)](/models/native-belebele-llama-1-13b-5-shot) - Score: not computed
- [Llama 1 30B (Belebele 5-shot)](/models/native-belebele-llama-1-30b-5-shot) - Score: not computed
- [Llama 1 65B (Belebele 5-shot)](/models/native-belebele-llama-1-65b-5-shot) - Score: not computed
- [Llama 1 7B (Belebele 5-shot)](/models/native-belebele-llama-1-7b-5-shot) - Score: not computed
- [Llama 2 base 70B (Belebele 5-shot)](/models/native-belebele-llama-2-base-70b-5-shot) - Score: not computed
- [Llama-2-chat 70B (Belebele 89-language English instructions)](/models/native-belebele-llama-2-chat-70b-89-language-english-instructions) - Score: not computed
- [Llama-2-chat 70B (Belebele 89-language translated instructions)](/models/native-belebele-llama-2-chat-70b-89-language-translated-instructions) - Score: not computed
- [Llama-2-chat 70B (Belebele 91-language in-language)](/models/native-belebele-llama-2-chat-70b-91-language-in-language) - Score: not computed
- [Llama-2-chat 70B (Belebele 91-language Translate-Test)](/models/native-belebele-llama-2-chat-70b-91-language-translate-test) - Score: not computed
- [Llama-2-chat 70B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-70b-zero-shot) - Score: not computed
- [Llama-2-chat 7B (Belebele zero-shot)](/models/native-belebele-llama-2-chat-7b-zero-shot) - Score: not computed
- [XLM-R large (550M) (Belebele English fine-tuning)](/models/native-belebele-xlm-r-large-550m-english-fine-tuning) - Score: not computed
- [XLM-R large (550M) (Belebele Translate-Train-All)](/models/native-belebele-xlm-r-large-550m-translate-train-all) - Score: not computed
- [XLM-V large (1.2B) (Belebele English fine-tuning)](/models/native-belebele-xlm-v-large-1-2b-english-fine-tuning) - Score: not computed
- [XLM-V large (1.2B) (Belebele Translate-Train-All)](/models/native-belebele-xlm-v-large-1-2b-translate-train-all) - Score: not computed
