Skip to main content
BenchLM

Belebele

We show this table for reference; we do not rank on it.

Belebele tests passage comprehension through questions with four answer options. Its parallel language variants let researchers study how reading performance changes across languages.

Original benchmark results

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

7 source tables · 20 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)

Paper summary across 122 language variants

Source vocabulary and parameter sizes are metadata, not scores. Source summary averages remain unchanged despite appendix discrepancies.

Published source

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper summary across 122 language variants. Published source metrics with original precision; unreported cells are not zero.
ModelSize or variantVocabulary sizeAverage accuracy (%)Language share ≥50% accuracy (%)Language share ≥70% accuracy (%)English accuracy (%)Non-English average (%)
5-Shot In-Context Learning (examples in English)
Llama 17B32K27.70.0%0.0%37.327.6
Llama 113B32K30.40.8%0.0%53.330.2
Llama 130B32K36.218.0%0.8%73.135.9
Llama 170B32K40.925.4%12.3%82.540.5
Llama 2 base70B32K48.038.5%26.2%90.947.7
Falcon40B65K37.316.4%1.6%77.236.9
Zero-Shot for Instructed Models (English instructions)
BLOOMZ**7.1B251K43.228.7%9.0%79.642.9
Llama-2-chat7B32K34.44.1%0.0%58.634.1
Llama-2-chat70B32K41.527.0%2.5%78.841.2
GPT3.5-turbounk100K51.144.2%29.2%87.750.7
Full Finetuning in English
XLM-Rlarge (550M)250K54.064.8%15.6%76.253.8
XLM-Vlarge (1.2B)902K55.669.7%21.2%76.254.9
InfoXLMlarge (550M)250K56.267.2%28.7%79.356.0
Translate-Train-All
XLM-Rlarge (550M)250K58.969.7%36.1%78.758.8
XLM-Vlarge (1.2B)902K60.276.2%32.8%77.860.1
InfoXLMlarge (550M)250K60.070.5%36.9%81.259.8

All 122 language results: multilingual encoders

Accuracy values are percentages. PCT rows are the percentage of language variants meeting an accuracy threshold; summary and appendix PCT values differ and are both retained.

Published source

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

All 122 language results: multilingual encoders. Published source metrics with original precision; unreported cells are not zero.
Language or summary metricXLM-V large (1.2B) (Belebele English fine-tuning)InfoXLM large (550M) (Belebele English fine-tuning)XLM-R large (550M) (Belebele English fine-tuning)XLM-V large (1.2B) (Belebele Translate-Train-All)InfoXLM large (550M) (Belebele Translate-Train-All)XLM-R large (550M) (Belebele Translate-Train-All)
AVG55.656.254.060.260.058.9
PCT Above 5069.7%67.2%64.8%76.2%70.5%69.7%
PCT Above 7021.9%28.9%15.7%33.1%37.2%36.4%
eng_Latn76.279.376.277.881.278.7
acm_Arab51.257.355.455.357.659.2
afr_Latn69.372.769.172.375.174.3
als_Latn68.468.964.970.872.271.4
amh_Ethi53.152.952.661.660.060.7
apc_Arab56.158.857.957.760.661.9
arb_Arab67.271.069.870.675.074.3
arb_Latn29.332.227.631.633.430.6
ars_Arab55.659.958.961.165.865.9
ary_Arab43.848.744.048.052.852.6
arz_Arab56.960.257.661.464.966.1
asm_Beng53.753.649.358.658.856.9
azj_Latn59.761.359.065.065.665.1
bam_Latn34.234.933.239.239.136.9
ben_Beng60.063.459.665.669.663.7
ben_Latn46.836.938.853.042.748.1
bod_Tibt24.024.923.724.823.336.9
bul_Cyrl72.672.070.174.075.374.2
cat_Latn71.674.472.075.778.174.7
ceb_Latn45.444.142.352.052.650.7
ces_Latn69.972.369.972.376.274.4
ckb_Arab29.752.330.336.958.036.9
dan_Latn70.874.172.973.076.374.7
deu_Latn72.675.772.974.178.776.7
ell_Grek70.372.370.373.174.973.0
est_Latn63.267.264.868.770.770.4
eus_Latn63.666.164.868.270.870.3
fin_Latn69.172.472.273.075.274.9
fra_Latn73.174.272.174.676.875.6
fuv_Latn29.727.726.432.830.731.1
gaz_Latn48.833.836.452.636.043.3
grn_Latn53.937.837.959.640.641.9
guj_Gujr58.757.054.163.365.963.1
hat_Latn57.139.635.263.244.139.8
hau_Latn51.041.148.253.448.153.0
heb_Hebr67.268.264.869.372.370.6
hin_Deva57.960.257.463.864.363.4
hin_Latn53.149.746.857.655.458.9
hrv_Latn70.072.469.971.275.374.0
hun_Latn69.770.870.073.174.272.8
hye_Armn59.461.058.965.966.164.7
ibo_Latn40.132.231.246.832.032.2
ilo_Latn37.436.333.838.140.639.7
ind_Latn68.970.768.071.373.170.4
isl_Latn67.366.063.870.168.969.0
ita_Latn70.672.870.071.876.473.3
jav_Latn64.259.860.867.263.366.8
jpn_Jpan66.470.167.671.371.871.0
kac_Latn32.029.132.133.834.033.3
kan_Knda61.162.059.766.668.469.1
kat_Geor64.764.863.668.068.967.4
kaz_Cyrl60.161.656.864.965.364.7
kea_Latn44.045.244.948.747.748.1
khk_Cyrl56.758.857.861.164.664.2
khm_Khmr60.059.057.763.064.263.8
kin_Latn35.933.634.339.139.138.6
kir_Cyrl65.463.461.868.368.267.7
kor_Hang70.171.468.772.974.674.8
lao_Laoo55.857.653.063.263.663.0
lin_Latn44.733.230.650.935.334.4
lit_Latn68.369.467.271.772.972.0
lug_Latn39.929.431.647.834.734.7
luo_Latn30.330.930.833.734.933.2
lvs_Latn70.171.368.774.175.673.0
mal_Mlym62.065.062.769.168.367.1
mar_Deva62.665.260.869.268.867.2
mkd_Cyrl67.869.365.771.073.872.8
mlt_Latn37.957.138.140.263.742.7
mri_Latn32.030.632.233.035.734.0
mya_Mymr56.659.153.662.265.162.9
nld_Latn68.471.771.068.674.072.8
nob_Latn71.873.670.772.875.474.2
npi_Deva58.460.755.764.465.862.7
npi_Latn38.335.833.837.436.434.8
nso_Latn45.931.330.053.234.134.7
nya_Latn31.029.229.834.233.030.8
ory_Orya60.862.158.665.665.463.9
pan_Guru58.159.257.863.162.662.0
pbt_Arab55.456.051.060.662.661.1
pes_Arab68.369.168.270.873.672.0
plt_Latn55.745.652.761.753.458.1
pol_Latn69.070.467.472.173.772.7
por_Latn70.974.370.673.877.174.0
ron_Latn72.372.971.374.076.274.8
rus_Cyrl71.973.872.275.476.877.1
shn_Mymr26.925.226.325.026.427.0
sin_Latn24.934.230.741.738.337.3
sin_Sinh64.467.262.769.870.268.6
slk_Latn69.371.970.272.676.773.0
slv_Latn69.772.268.671.875.473.9
sna_Latn34.837.233.237.138.635.9
snd_Arab55.256.651.960.061.361.3
som_Latn46.039.142.650.746.350.7
sot_Latn46.829.331.352.031.932.7
spa_Latn71.073.371.472.775.376.4
srp_Cyrl71.070.971.173.676.175.9
ssw_Latn39.830.634.347.134.338.9
sun_Latn60.950.755.364.255.859.4
swe_Latn73.075.074.274.276.975.1
swh_Latn64.965.362.869.369.268.7
tam_Taml61.864.661.767.469.465.3
tel_Telu55.657.853.662.163.261.1
tgk_Cyrl38.258.633.839.264.339.6
tgl_Latn69.267.464.772.070.470.0
tha_Thai63.868.165.869.068.970.1
tir_Ethi33.336.733.839.942.137.7
tsn_Latn49.035.030.849.835.734.3
tso_Latn37.936.334.241.739.737.1
tur_Latn66.770.266.870.672.072.0
ukr_Cyrl70.470.971.072.374.975.0
urd_Arab61.663.859.365.668.666.3
urd_Latn42.242.640.849.448.948.4
uzn_Latn65.266.964.469.170.670.2
vie_Latn69.671.169.473.772.971.4
war_Latn46.444.743.747.649.346.6
wol_Latn36.832.230.440.632.332.2
xho_Latn48.736.139.054.440.245.4
yor_Latn35.029.328.738.632.027.9
zho_Hans69.874.671.073.776.274.8
zho_Hant69.272.467.173.174.371.3
zsm_Latn69.172.669.972.473.372.2
zul_Latn46.936.439.054.239.844.1

Multilingual encoder training settings

These are training settings, not benchmark scores. XLM-R Translate-Train-All learning rate is printed as 3-6 in the source; no exponent is inferred.

Published source

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Multilingual encoder training settings. Published source metrics with original precision; unreported cells are not zero.
Training settingXLM-V large (1.2B) (Belebele English fine-tuning)InfoXLM large (550M) (Belebele English fine-tuning)XLM-R large (550M) (Belebele English fine-tuning)XLM-V large (1.2B) (Belebele Translate-Train-All)InfoXLM large (550M) (Belebele Translate-Train-All)XLM-R large (550M) (Belebele Translate-Train-All)
epochs343111
training set size67.5k67.5k67.5k650k650k650k
learning rate5e-64e-65e-63-63e-63e-6
weight decay0.010.010.010.0010.0010.001
batch size646464646464

All 122 language results: large language models

The GPT-3.5 appendix average is 50.6; the summary prints 51.1. The Llama 1 checkpoint is labeled 65B in this appendix.

Published source

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

All 122 language results: large language models. Published source metrics with original precision; unreported cells are not zero.
Language or summary metricGPT3.5-turbo unk (Belebele zero-shot)Llama-2-chat 70B (Belebele zero-shot)Llama 2 base 70B (Belebele 5-shot)Llama 1 65B (Belebele 5-shot)Falcon 40B (Belebele 5-shot)XLM-V large (1.2B) (Belebele Translate-Train-All)
AVG50.641.548.040.937.360.2
PCT Above 5043.4 %27.1%38.5 %25.4%16.4%76.2%
PCT Above 7028.9 %2.5%26.2 %12.3%1.6%33.1%
eng_Latn87.778.890.982.577.277.8
acm_Arab51.635.947.937.937.655.3
afr_Latn78.357.975.960.753.472.3
als_Latn67.136.045.434.936.670.8
amh_Ethi28.728.927.527.824.861.6
apc_Arab55.638.851.239.636.357.7
arb_Arab69.342.361.744.138.370.6
arb_Latn31.130.226.828.026.331.6
ars_Arab55.137.450.240.732.161.1
ary_Arab45.732.640.633.132.348.0
arz_Arab56.737.350.737.433.061.4
asm_Beng36.035.732.328.922.458.6
azj_Latn54.933.442.233.634.165.0
bam_Latn31.729.430.328.429.739.2
ben_Beng43.634.939.133.422.665.6
ben_Latn34.630.429.629.232.153.0
bod_Tibt26.628.325.724.926.824.8
bul_Cyrl76.065.080.469.341.974.0
cat_Latn78.468.284.676.358.875.7
ceb_Latn53.340.650.438.939.252.0
ces_Latn76.965.081.170.765.072.3
ckb_Arab31.832.828.731.628.936.9
dan_Latn80.766.283.673.656.273.0
deu_Latn83.369.484.676.070.174.1
ell_Grek73.050.764.944.231.273.1
est_Latn73.136.653.036.334.968.7
eus_Latn40.931.134.732.838.968.2
fin_Latn77.962.779.355.742.873.0
fra_Latn83.172.286.477.569.774.6
fuv_Latn26.129.824.925.425.132.8
gaz_Latn30.329.327.829.124.952.6
grn_Latn34.232.232.430.333.859.6
guj_Gujr38.431.127.125.724.763.3
hat_Latn51.634.137.433.736.263.2
hau_Latn32.232.128.026.428.953.4
heb_Hebr64.241.454.941.431.169.3
hin_Deva49.142.052.638.427.163.8
hin_Latn52.339.249.034.240.057.6
hrv_Latn78.464.779.866.948.771.2
hun_Latn74.661.178.866.737.773.1
hye_Armn35.031.934.132.125.465.9
ibo_Latn28.430.127.425.330.246.8
ilo_Latn37.133.236.632.135.138.1
ind_Latn74.261.381.455.752.171.3
isl_Latn62.338.054.342.136.470.1
ita_Latn80.068.684.576.166.471.8
jav_Latn46.737.040.333.036.867.2
jpn_Jpan70.956.677.653.949.671.3
kac_Latn30.930.727.728.627.833.8
kan_Knda40.621.925.724.424.066.6
kat_Geor33.034.637.834.323.468.0
kaz_Cyrl35.032.429.332.432.664.9
kea_Latn46.038.145.438.138.048.7
khk_Cyrl32.031.129.828.427.461.1
khm_Khmr30.430.627.028.225.063.0
kin_Latn35.230.629.828.531.939.1
kir_Cyrl37.932.234.632.531.968.3
kor_Hang67.156.377.852.940.272.9
lao_Laoo30.026.524.326.228.163.2
lin_Latn33.831.028.030.429.350.9
lit_Latn72.039.752.139.639.371.7
lug_Latn28.430.929.228.328.947.8
luo_Latn27.131.229.429.329.933.7
lvs_Latn70.841.051.339.037.674.1
mal_Mlym34.930.132.430.021.269.1
mar_Deva38.334.841.232.925.069.2
mkd_Cyrl69.455.772.556.238.171.0
mlt_Latn44.836.244.936.735.440.2
mri_Latn33.331.828.532.029.733.0
mya_Mymr30.331.324.124.222.662.2
nld_Latn80.466.282.273.366.768.6
nob_Latn79.065.781.870.960.872.8
npi_Deva40.432.940.433.025.464.4
npi_Latn35.130.430.230.030.937.4
nso_Latn33.630.130.427.429.353.2
nya_Latn33.229.327.328.729.334.2
ory_OryaUnreported29.224.823.923.765.6
pan_Guru39.133.126.327.123.463.1
pbt_Arab32.330.230.829.429.460.6
pes_Arab61.841.853.941.035.970.8
plt_Latn32.330.529.631.031.461.7
pol_Latn74.761.779.267.059.972.1
por_Latn83.070.286.175.468.373.8
ron_Latn77.465.683.473.266.674.0
rus_Cyrl78.467.082.773.148.175.4
shn_MymrUnreported28.225.622.724.025.0
sin_Latn30.431.933.827.932.641.7
sin_Sinh32.633.425.229.427.769.8
slk_Latn77.358.875.260.457.072.6
slv_Latn77.462.476.765.643.771.8
sna_Latn35.430.227.428.331.637.1
snd_Arab34.129.730.928.930.260.0
som_Latn32.430.327.827.629.950.7
sot_Latn33.930.028.926.829.952.0
spa_Latn79.268.485.074.869.272.7
srp_Cyrl74.865.181.070.740.273.6
ssw_Latn32.030.727.728.030.147.1
sun_Latn38.934.937.830.734.164.2
swe_Latn81.767.482.773.767.374.2
swh_Latn70.335.139.634.436.769.3
tam_Taml32.834.433.231.624.467.4
tel_Telu34.627.525.926.622.462.1
tgk_Cyrl37.732.534.033.132.739.2
tgl_Latn66.749.668.148.347.772.0
tha_Thai55.738.946.235.033.069.0
tir_Ethi28.429.624.523.525.039.9
tsn_Latn31.830.128.524.731.249.8
tso_Latn33.430.030.428.029.741.7
tur_Latn69.947.365.442.139.670.6
ukr_Cyrl72.865.780.869.741.972.3
urd_Arab48.337.043.234.731.765.6
urd_Latn40.334.138.030.134.249.4
uzn_Latn44.133.135.130.633.169.1
vie_Latn72.959.678.443.541.473.7
war_Latn48.939.344.437.438.647.6
wol_Latn29.028.927.626.026.840.6
xho_Latn30.029.928.227.630.254.4
yor_Latn29.130.128.327.727.238.6
zho_Hans77.662.483.764.666.073.7
zho_Hant76.359.382.057.762.273.1
zsm_Latn74.056.476.351.751.372.4
zul_Latn30.430.229.727.130.754.2

Multiple-script language comparison

Per-language accuracy (%). The final AVG column is a cross-system summary, not an additional model.

Published source

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Multiple-script language comparison. Published source metrics with original precision; unreported cells are not zero.
LanguageGPT3.5-turbo unk (Belebele zero-shot)Llama 2 base 70B (Belebele 5-shot)Falcon 40B (Belebele 5-shot)InfoXLM large (550M) (Belebele English fine-tuning)Average across shown systems (%)
arb_Arab69.361.738.371.060.1
arb_Latn31.126.826.332.229.1
ben_Beng43.639.122.663.442.2
ben_Latn34.629.632.136.933.3
hin_Deva49.152.627.160.247.3
hin_Latn52.349.040.049.747.8
npi_Deva40.440.425.460.741.7
npi_Latn35.130.230.935.833.0
sin_Sinh32.625.227.767.238.2
sin_Latn30.433.832.634.232.8
urd_Arab48.343.231.763.846.7
urd_Latn40.338.034.242.638.8
zho_Hant76.382.062.272.473.3
zho_Hans77.683.766.074.675.5

All translated-test language results

The 91 non-English-language subset also shows an English reference row. Its averages are not the 122-language averages. In-language appendix average is 44.0; the summary prints 44.1.

Published source

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

All translated-test language results. Published source metrics with original precision; unreported cells are not zero.
Language or summary metricLlama-2-chat 70B (Belebele 91-language in-language)Llama-2-chat 70B (Belebele 91-language Translate-Test)XLM-V large (1.2B) (Belebele Translate-Train-All)
AVG44.057.164.9
PCT Above 5035.2%78.0%90.1%
PCT Above 702.2%2.2%42.9%
eng_Latn78.878.877.8
fra_Latn72.270.673.1
por_Latn70.269.970.9
deu_Latn69.465.772.6
ita_Latn68.666.170.6
spa_Latn68.469.371.0
cat_Latn68.267.071.6
swe_Latn67.466.173.0
rus_Cyrl67.067.371.9
dan_Latn66.266.870.8
nld_Latn66.267.268.4
nob_Latn65.768.371.8
ukr_Cyrl65.766.070.4
ron_Latn65.667.072.3
srp_Cyrl65.166.271.0
bul_Cyrl65.067.772.6
ces_Latn65.065.669.9
hrv_Latn64.765.370.0
fin_Latn62.761.169.1
slv_Latn62.461.269.7
zho_Hans62.471.269.8
pol_Latn61.763.069.0
ind_Latn61.364.868.9
hun_Latn61.162.969.7
vie_Latn59.659.469.6
zho_Hant59.365.869.2
slk_Latn58.866.269.3
afr_Latn57.965.069.3
jpn_Jpan56.654.866.4
zsm_Latn56.467.069.1
kor_Hang56.356.770.1
mkd_Cyrl55.766.767.8
ell_Grek50.767.670.3
tgl_Latn49.662.269.2
tur_Latn47.362.666.7
arb_Arab42.360.767.2
hin_Deva42.062.857.9
pes_Arab41.859.668.3
heb_Hebr41.462.067.2
lvs_Latn41.060.970.1
ceb_Latn40.662.645.4
lit_Latn39.760.868.3
hin_Latn39.252.753.1
tha_Thai38.954.163.8
isl_Latn38.058.167.3
jav_Latn37.055.364.2
urd_Arab37.059.461.6
est_Latn36.659.463.2
als_Latn36.063.168.4
asm_Beng35.757.753.7
swh_Latn35.157.864.9
ben_Beng34.961.060.0
sun_Latn34.950.860.9
mar_Deva34.860.062.6
kat_Geor34.657.764.7
tam_Taml34.455.961.8
urd_Latn34.143.042.2
hat_Latn34.156.357.1
azj_Latn33.455.659.7
sin_Sinh33.457.764.4
pan_Guru33.157.658.1
npi_Deva32.962.058.4
ckb_Arab32.851.329.7
kaz_Cyrl32.453.260.1
hau_Latn32.143.451.0
hye_Armn31.958.059.4
mya_Mymr31.346.656.6
khk_Cyrl31.152.256.7
guj_Gujr31.159.658.7
lin_Latn31.040.344.7
lug_Latn30.938.739.9
ssw_Latn30.743.239.8
khm_Khmr30.652.860.0
plt_Latn30.546.755.7
ben_Latn30.445.146.8
som_Latn30.340.846.0
pbt_Arab30.248.855.4
zul_Latn30.244.446.9
nso_Latn30.143.445.9
tsn_Latn30.140.449.0
yor_Latn30.137.735.0
ibo_Latn30.135.340.1
mal_Mlym30.163.062.0
xho_Latn29.949.248.7
fuv_Latn29.829.429.7
gaz_Latn29.337.048.8
ory_Orya29.257.860.8
amh_Ethi28.950.453.1
wol_Latn28.939.036.8
tel_Telu27.554.355.6
lao_Laoo26.547.455.8
kan_Knda21.962.061.1

Translation and instruction-language comparisons

The first two rows cover 91 non-English languages; the instruction-language comparison covers 89. These subsets are distinct from the full 122-language result.

Published source

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Translation and instruction-language comparisons. Published source metrics with original precision; unreported cells are not zero.
ModelVariantEvaluation settingAverage accuracy (%)Language share ≥50% accuracy (%)Language share ≥70% accuracy (%)English accuracy (%)
Translate-Test (English) on 91 non-English languages in Zero-Shot
Llama-2-chat70BTranslate-Test57.178.0%2.2%78.8
Llama-2-chat70BIn-Language44.135.2%2.2%78.8
Translated Instructions in 89 non-English languages Zero-Shot
Llama-2-chat70BIn-Language Translated Instructions38.736.0%7.9%78.8
Llama-2-chat70BEnglish Instructions44.937.1%3.4%78.8

Original benchmark results and configurations

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.

Snapshot

7 source tables20 configurationsDisplay only

About Belebele

Year

2023

Tasks

Multilingual multiple-choice reading comprehension

Format

Decision classification

Difficulty

Depends on the evaluated split and label mapping

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

Freshness and provenance

Version

Belebele 2023

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Which Belebele results are included?

This page includes 7 numeric result and configuration tables from Belebele’s original paper and available benchmark-owner updates, covering 20 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

Can I compare these Belebele scores with the Perplexity panel?

Compare Belebele scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

Where can I download the Belebele results?

The results download on this page provides every imported Belebele table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.

Last updated: October 1, 2026 source review · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.