# Arena-Hard (gemini-1.5-pro-001) (JudgeBench) Benchmark Scores & Performance

> Arena-Hard (gemini-1.5-pro-001) (JudgeBench) has published results in 3 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-judgebench-arena-hard-gemini-1-5-pro-001

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | JudgeBench authors |
| Source Type | Evaluated configuration |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://huggingface.co/spaces/ScalerLab/JudgeBench) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Arena-Hard (gemini-1.5-pro-001) (JudgeBench)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-judgebench-arena-hard-gemini-1-5-pro-001) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-judgebench-arena-hard-gemini-1-5-pro-001&format=csv)

### Maintainer results on GPT-4o response pairs

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

GPT-4o-2024-05-13 response split. The public app rounds to one decimal; paper tables retain two decimals.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://huggingface.co/spaces/ScalerLab/JudgeBench) · [Full JudgeBench results](/benchmarks/judgebench)

| Judge | Type | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- | --- |
| Arena-Hard (gemini-1.5-pro-001) | Prompted Judge | 49.4 | 42.9 | 64.3 | 26.2 | 47.1 |

### Paper: Arena-Hard prompted judges

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2) · [Full JudgeBench results](/benchmarks/judgebench)

| Judge model | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| Gemini-1.5-pro | 49.35 | 42.86 | 64.29 | 26.19 | 47.14 |

### Paper: solving versus judging

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

Paper v2; configurations and precision remain as published. Solver results are direct answers rather than judge preferences.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2) · [Full JudgeBench results](/benchmarks/judgebench)

| Setup | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| Gemini-1.5-pro Judge | 49.35 | 42.86 | 64.29 | 26.19 | 47.14 |

## Other JudgeBench authors Models

- [Arena-Hard (claude-3-5-sonnet-20240620) (JudgeBench)](/models/native-judgebench-arena-hard-claude-3-5-sonnet-20240620) - Score: not computed
- [Arena-Hard (claude-3-haiku-20240307) (JudgeBench)](/models/native-judgebench-arena-hard-claude-3-haiku-20240307) - Score: not computed
- [Arena-Hard (DeepSeek-R1-250120) (JudgeBench)](/models/native-judgebench-arena-hard-deepseek-r1-250120) - Score: not computed
- [Arena-Hard (gemini-1.5-flash-001) (JudgeBench)](/models/native-judgebench-arena-hard-gemini-1-5-flash-001) - Score: not computed
- [Arena-Hard (gpt-4o-2024-05-13) (JudgeBench)](/models/native-judgebench-arena-hard-gpt-4o-2024-05-13) - Score: not computed
- [Arena-Hard (gpt-4o-mini-2024-07-18) (JudgeBench)](/models/native-judgebench-arena-hard-gpt-4o-mini-2024-07-18) - Score: not computed
- [Arena-Hard (Llama-3.1-405B-Instruct) (JudgeBench)](/models/native-judgebench-arena-hard-llama-3-1-405b-instruct) - Score: not computed
- [Arena-Hard (Llama-3.1-70B-Instruct) (JudgeBench)](/models/native-judgebench-arena-hard-llama-3-1-70b-instruct) - Score: not computed
- [Arena-Hard (Llama-3.1-8B-Instruct) (JudgeBench)](/models/native-judgebench-arena-hard-llama-3-1-8b-instruct) - Score: not computed
- [Arena-Hard (o1-mini-2024-09-12) (JudgeBench)](/models/native-judgebench-arena-hard-o1-mini-2024-09-12) - Score: not computed
- [Arena-Hard (o1-preview-2024-09-12) (JudgeBench)](/models/native-judgebench-arena-hard-o1-preview-2024-09-12) - Score: not computed
- [Arena-Hard (o3-mini-2025-01-31 (high)) (JudgeBench)](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-high) - Score: not computed
- [Arena-Hard (o3-mini-2025-01-31 (low)) (JudgeBench)](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-low) - Score: not computed
- [Arena-Hard (o3-mini-2025-01-31 (medium)) (JudgeBench)](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-medium) - Score: not computed
- [ChatEval (gpt-4o-2024-05-13) (JudgeBench)](/models/native-judgebench-chateval-gpt-4o-2024-05-13) - Score: not computed
- [Claude-3.5-Sonnet Solver (JudgeBench)](/models/native-judgebench-claude-3-5-sonnet-solver) - Score: not computed
- [Gemini-1.5-pro Solver (JudgeBench)](/models/native-judgebench-gemini-1-5-pro-solver) - Score: not computed
- [GPT-4o Solver (JudgeBench)](/models/native-judgebench-gpt-4o-solver) - Score: not computed
- [Llama-3.1-405B-Instruct Solver (JudgeBench)](/models/native-judgebench-llama-3-1-405b-instruct-solver) - Score: not computed
- [Vanilla (gpt-4o-2024-05-13) (JudgeBench)](/models/native-judgebench-vanilla-gpt-4o-2024-05-13) - Score: not computed
- [VertexAI Evaluation (gemini-1.5-pro-001) (JudgeBench)](/models/native-judgebench-vertexai-evaluation-gemini-1-5-pro-001) - Score: not computed
