Skip to main content
BenchLM

Data as of October 1, 2026 · How the score is built

TrackedProprietaryUnspecified

Benchmark-owner results and configuration

Llama 2 base 70B (Belebele 5-shot)

Decision readingLlama 2 base 70B (Belebele 5-shot) is tracked, but not publicly ranked yet. The original benchmark tables below include the published metrics for this evaluated configuration. Their distinct protocols remain outside general model rankings.

Original benchmark results

Every published metric for this configuration, with source precision and evaluation limits. Download results (JSON) · Numeric metrics (CSV)

Paper summary across 122 language variants

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

Source vocabulary and parameter sizes are metadata, not scores. Source summary averages remain unchanged despite appendix discrepancies.

Published source · Full Belebele results

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper summary across 122 language variants. Published source metrics with original precision; unreported cells are not zero.
ModelSize or variantVocabulary sizeAverage accuracy (%)Language share ≥50% accuracy (%)Language share ≥70% accuracy (%)English accuracy (%)Non-English average (%)
Llama 2 base70B32K48.038.5%26.2%90.947.7

All 122 language results: large language models

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

The GPT-3.5 appendix average is 50.6; the summary prints 51.1. The Llama 1 checkpoint is labeled 65B in this appendix.

Published source · Full Belebele results

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

All 122 language results: large language models. Published source metrics with original precision; unreported cells are not zero.
Language or summary metricLlama 2 base 70B (Belebele 5-shot)
AVG48.0
PCT Above 5038.5 %
PCT Above 7026.2 %
eng_Latn90.9
acm_Arab47.9
afr_Latn75.9
als_Latn45.4
amh_Ethi27.5
apc_Arab51.2
arb_Arab61.7
arb_Latn26.8
ars_Arab50.2
ary_Arab40.6
arz_Arab50.7
asm_Beng32.3
azj_Latn42.2
bam_Latn30.3
ben_Beng39.1
ben_Latn29.6
bod_Tibt25.7
bul_Cyrl80.4
cat_Latn84.6
ceb_Latn50.4
ces_Latn81.1
ckb_Arab28.7
dan_Latn83.6
deu_Latn84.6
ell_Grek64.9
est_Latn53.0
eus_Latn34.7
fin_Latn79.3
fra_Latn86.4
fuv_Latn24.9
gaz_Latn27.8
grn_Latn32.4
guj_Gujr27.1
hat_Latn37.4
hau_Latn28.0
heb_Hebr54.9
hin_Deva52.6
hin_Latn49.0
hrv_Latn79.8
hun_Latn78.8
hye_Armn34.1
ibo_Latn27.4
ilo_Latn36.6
ind_Latn81.4
isl_Latn54.3
ita_Latn84.5
jav_Latn40.3
jpn_Jpan77.6
kac_Latn27.7
kan_Knda25.7
kat_Geor37.8
kaz_Cyrl29.3
kea_Latn45.4
khk_Cyrl29.8
khm_Khmr27.0
kin_Latn29.8
kir_Cyrl34.6
kor_Hang77.8
lao_Laoo24.3
lin_Latn28.0
lit_Latn52.1
lug_Latn29.2
luo_Latn29.4
lvs_Latn51.3
mal_Mlym32.4
mar_Deva41.2
mkd_Cyrl72.5
mlt_Latn44.9
mri_Latn28.5
mya_Mymr24.1
nld_Latn82.2
nob_Latn81.8
npi_Deva40.4
npi_Latn30.2
nso_Latn30.4
nya_Latn27.3
ory_Orya24.8
pan_Guru26.3
pbt_Arab30.8
pes_Arab53.9
plt_Latn29.6
pol_Latn79.2
por_Latn86.1
ron_Latn83.4
rus_Cyrl82.7
shn_Mymr25.6
sin_Latn33.8
sin_Sinh25.2
slk_Latn75.2
slv_Latn76.7
sna_Latn27.4
snd_Arab30.9
som_Latn27.8
sot_Latn28.9
spa_Latn85.0
srp_Cyrl81.0
ssw_Latn27.7
sun_Latn37.8
swe_Latn82.7
swh_Latn39.6
tam_Taml33.2
tel_Telu25.9
tgk_Cyrl34.0
tgl_Latn68.1
tha_Thai46.2
tir_Ethi24.5
tsn_Latn28.5
tso_Latn30.4
tur_Latn65.4
ukr_Cyrl80.8
urd_Arab43.2
urd_Latn38.0
uzn_Latn35.1
vie_Latn78.4
war_Latn44.4
wol_Latn27.6
xho_Latn28.2
yor_Latn28.3
zho_Hans83.7
zho_Hant82.0
zsm_Latn76.3
zul_Latn29.7

Multiple-script language comparison

Belebele reports multiple-choice accuracy over 122 language variants, with 900 questions per variant. Few-shot, zero-shot, English fine-tuning, translated training, and translated tests keep separate configurations. All published per-language values are included. Summary and appendix averages sometimes disagree: GPT-3.5 is 51.1 versus 50.6, and the 91-language Llama-2-chat subset is 44.1 versus 44.0. Table 2 labels a Llama 1 row 70B while Table 7 identifies 65B; the profile uses the appendix identity and retains the conflicting label. BLOOMZ used FLORES material during training, which can advantage its evaluation.

Per-language accuracy (%). The final AVG column is a cross-system summary, not an additional model.

Published source · Full Belebele results

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants — Lucas Bandarkar and colleagues. CC BY SA 4.0. Copyright 2024 the authors. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Multiple-script language comparison. Published source metrics with original precision; unreported cells are not zero.
LanguageLlama 2 base 70B (Belebele 5-shot)
arb_Arab61.7
arb_Latn26.8
ben_Beng39.1
ben_Latn29.6
hin_Deva52.6
hin_Latn49.0
npi_Deva40.4
npi_Latn30.2
sin_Sinh25.2
sin_Latn33.8
urd_Arab43.2
urd_Latn38.0
zho_Hant82.0
zho_Hans83.7

Lineage

The sequence follows explicit supersedes links. Each score is estimated for that model; a relative can inform a sparse estimate but never sets a floor, so a newer release can score below an earlier one. Scores and prices remain blank when the corresponding public row or first-party rate is unavailable.

  1. Release date not sourced · you are here

    Llama 2 base 70B (Belebele 5-shot)

    Not publicly ranked · Price not listed

Benchmark-system

Spec sheet

Each documented value carries its source. Missing fields stay visible as not sourced or not published, rather than disappearing from the page.

API model ID
Not publishedProvider pricing
Context window
Not published
Maximum output
Not sourced yet
Knowledge cutoff
Not sourced yet
Input modalities
Not sourced yet
Output modalities
Not sourced yet
Parameters
Not sourced yet
Availability
Not sourced yet
Cloud regions
Not tracked yet
Lifecycle
Tracked
API capabilities
Tool calling, structured outputs, and batch support are not tracked yet
Prompt caching
Not documented in the pricing recordProvider pricing
Self-host
Weights are not published
Rate limits
Not tracked yet

How to read this profile

The visual layer above carries the decisions. These notes preserve the model, ranking, coverage, and family context behind the numbers.

The original benchmark tables show this evaluated configuration’s published results. Training choices, prompts, response splits, and metrics remain attached to each table. The profile stays outside general model rankings.

Llama 2 base 70B (Belebele 5-shot) is a research system model. No explicit reasoning mode is documented in this profile.

Evaluated system: Llama 2 base 70B (Belebele 5-shot). The linked benchmark-owner report supplies the full numeric results and protocol. This profile represents that evaluated configuration; it does not transfer results to a base model or another prompt. Context limit, current API tariff, and release date are not established by this evaluation.

The original benchmark results appear in the source tables above; they remain separate from weighted benchmark slots.

Last updated October 1, 2026. Runtime fields remain blank until a sourced snapshot exists.

Questions

How does Llama 2 base 70B (Belebele 5-shot) perform overall in AI benchmarks?

This configuration has original benchmark results in the source tables on this profile. Every published metric retains its evaluated configuration, source precision, and protocol. These results stay outside general model rankings; they do not produce a public overall score or transfer to a base model with different settings.

Does Llama 2 base 70B (Belebele 5-shot) have full benchmark coverage on BenchLM?

This configuration has results in the original benchmark tables on this profile. Those tables preserve every imported source metric, but they do not fill the site’s general scoring slots. Other benchmarks remain unmeasured, and compatible configurations are required before comparing results across different reports.

What is the context window size of Llama 2 base 70B (Belebele 5-shot)?

Llama 2 base 70B (Belebele 5-shot)'s context window is not documented in a source tied to this exact model yet. The profile leaves the value unavailable instead of borrowing a limit from an earlier family member or an unverified route. Maximum output length remains a separate field.

Watch Llama 2 base 70B (Belebele 5-shot) in the weekly brief

Get one weekly email when material rank, price, availability, or benchmark evidence changes are worth revisiting.

Read a sample issue

Join 2,000+ readers.

Compare Llama 2 base 70B (Belebele 5-shot) with every tracked model782 comparisons