# JudgeLM-7B-v1.0 (JudgeBench) Benchmark Scores & Performance

> JudgeLM-7B-v1.0 (JudgeBench) has published results in 4 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-judgebench-judgelm-7b-v1-0

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | BAAI |
| Source Type | Open Weight |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Checkpoint identified by JudgeBench](https://huggingface.co/BAAI/JudgeLM-7B-v1.0) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: JudgeLM-7B-v1.0 (JudgeBench)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-judgebench-judgelm-7b-v1-0) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-judgebench-judgelm-7b-v1-0&format=csv)

### Maintainer results on GPT-4o response pairs

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

GPT-4o-2024-05-13 response split. The public app rounds to one decimal; paper tables retain two decimals.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://huggingface.co/spaces/ScalerLab/JudgeBench) · [Full JudgeBench results](/benchmarks/judgebench)

| Judge | Type | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- | --- |
| JudgeLM-7B-v1.0 | Fine-Tuned Judge | 23.4 | 29.6 | 32.1 | 11.9 | 25.1 |

### Paper: judge methods

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2) · [Full JudgeBench results](/benchmarks/judgebench)

| Judge | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| JudgeLM-7B | 23.38 | 29.59 | 32.14 | 11.90 | 25.14 |

### Paper: judgment outcomes

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2) · [Full JudgeBench results](/benchmarks/judgebench)

| Judge | A>B count | A<B count | Tie count | Invalid count |
| --- | --- | --- | --- | --- |
| JudgeLM-7B | 399 | 229 | 72 | 0 |

### Paper: order inconsistency

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2) · [Full JudgeBench results](/benchmarks/judgebench)

| Judge | Inconsistent (%) |
| --- | --- |
| JudgeLM-7B | 59.71% |

## Other BAAI Models

- [BAAI bge-reranker-v2-m3](/models/bge-reranker-v2-m3) - Score: not computed
- [JudgeLM-13B-v1.0 (JudgeBench)](/models/native-judgebench-judgelm-13b-v1-0) - Score: not computed
- [JudgeLM-33B-v1.0 (JudgeBench)](/models/native-judgebench-judgelm-33b-v1-0) - Score: not computed
