# JudgeBench

> JudgeBench evaluates whether a judge can identify the better response in pairs with objective correctness labels. It tests judgment on challenging questions rather than agreement with writing style preferences.

Canonical page: https://benchlm.ai/benchmarks/judgebench

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026 source review

## About JudgeBench

- Year: 2024
- Tasks: Choosing the better of two model responses
- Format: Decision classification
- Difficulty: Depends on the evaluated split and label mapping
- Paper: [JudgeBench: A Benchmark for Evaluating LLM-Based Judges](https://github.com/ScalerLab/JudgeBench)

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

JudgeBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Original benchmark results

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

[Download full results (JSON)](/api/data/decision-benchmarks?benchmark=judgeBench)

### Maintainer results on GPT-4o response pairs

GPT-4o-2024-05-13 response split. The public app rounds to one decimal; paper tables retain two decimals.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://huggingface.co/spaces/ScalerLab/JudgeBench)

| Judge | Type | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- | --- |
| [Arena-Hard (o3-mini-2025-01-31 (high))](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-high) | Prompted Judge | 67.5 | 89.8 | 87.5 | 100.0 | 80.9 |
| [Arena-Hard (o3-mini-2025-01-31 (medium))](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-medium) | Prompted Judge | 62.3 | 86.7 | 85.7 | 92.9 | 76.6 |
| [Arena-Hard (o1-preview-2024-09-12)](/models/native-judgebench-arena-hard-o1-preview-2024-09-12) | Prompted Judge | 66.2 | 79.6 | 85.7 | 85.7 | 75.4 |
| [Arena-Hard (DeepSeek-R1-250120)](/models/native-judgebench-arena-hard-deepseek-r1-250120) | Prompted Judge | 59.1 | 82.7 | 80.4 | 92.9 | 73.1 |
| [Arena-Hard (o3-mini-2025-01-31 (low))](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-low) | Prompted Judge | 63.0 | 69.4 | 83.9 | 83.3 | 70.6 |
| [Arena-Hard (o1-mini-2024-09-12)](/models/native-judgebench-arena-hard-o1-mini-2024-09-12) | Prompted Judge | 58.4 | 62.2 | 82.1 | 78.6 | 65.7 |
| [Arena-Hard (claude-3-5-sonnet-20240620)](/models/native-judgebench-arena-hard-claude-3-5-sonnet-20240620) | Prompted Judge | 62.3 | 66.3 | 66.1 | 64.3 | 64.3 |
| [Skywork-Reward-Gemma-2-27B](/models/native-judgebench-skywork-reward-gemma-2-27b) | Reward Model | 59.7 | 66.3 | 83.9 | 50.0 | 64.3 |
| [InternLM2-20B-Reward](/models/native-judgebench-internlm2-20b-reward) | Reward Model | 62.3 | 69.4 | 66.1 | 50.0 | 63.4 |
| [CompassJudger-1-32B](/models/native-judgebench-compassjudger-1-32b) | Fine-Tuned Judge | 58.4 | 62.2 | 75.0 | 59.5 | 62.3 |
| [Skywork-Reward-Llama-3.1-8B](/models/native-judgebench-skywork-reward-llama-3-1-8b) | Reward Model | 59.1 | 64.3 | 76.8 | 50.0 | 62.3 |
| [GRM-Gemma-2B](/models/native-judgebench-grm-gemma-2b) | Reward Model | 63.0 | 53.1 | 64.3 | 54.8 | 59.4 |
| [InternLM2-7B-Reward](/models/native-judgebench-internlm2-7b-reward) | Reward Model | 56.5 | 61.2 | 71.4 | 50.0 | 59.4 |
| [Skywork-Critic-Llama-3.1-70B](/models/native-judgebench-skywork-critic-llama-3-1-70b) | Fine-Tuned Judge | 55.8 | 55.1 | 73.2 | 47.6 | 57.4 |
| [Arena-Hard (Llama-3.1-405B-Instruct)](/models/native-judgebench-arena-hard-llama-3-1-405b-instruct) | Prompted Judge | 55.8 | 54.1 | 69.6 | 50.0 | 56.9 |
| [Arena-Hard (gpt-4o-2024-05-13)](/models/native-judgebench-arena-hard-gpt-4o-2024-05-13) | Prompted Judge | 50.6 | 54.1 | 75.0 | 59.5 | 56.6 |
| [CompassJudger-1-14B](/models/native-judgebench-compassjudger-1-14b) | Fine-Tuned Judge | 51.9 | 60.2 | 67.9 | 42.9 | 55.7 |
| [Skywork-Critic-Llama-3.1-8B](/models/native-judgebench-skywork-critic-llama-3-1-8b) | Fine-Tuned Judge | 51.3 | 54.1 | 73.2 | 33.3 | 53.4 |
| [Arena-Hard (Llama-3.1-70B-Instruct)](/models/native-judgebench-arena-hard-llama-3-1-70b-instruct) | Prompted Judge | 51.3 | 49.0 | 60.7 | 52.4 | 52.3 |
| [Vanilla (gpt-4o-2024-05-13)](/models/native-judgebench-vanilla-gpt-4o-2024-05-13) | Prompted Judge | 44.2 | 48.0 | 66.1 | 61.9 | 50.9 |
| [Arena-Hard (gpt-4o-mini-2024-07-18)](/models/native-judgebench-arena-hard-gpt-4o-mini-2024-07-18) | Prompted Judge | 48.1 | 43.9 | 69.6 | 45.2 | 50.0 |
| [Arena-Hard (gemini-1.5-pro-001)](/models/native-judgebench-arena-hard-gemini-1-5-pro-001) | Prompted Judge | 49.4 | 42.9 | 64.3 | 26.2 | 47.1 |
| [CompassJudger-1-7B](/models/native-judgebench-compassjudger-1-7b) | Fine-Tuned Judge | 42.2 | 37.8 | 69.6 | 47.6 | 46.0 |
| [CompassJudger-1-1.5B](/models/native-judgebench-compassjudger-1-1-5b) | Fine-Tuned Judge | 48.1 | 39.8 | 50.0 | 35.7 | 44.6 |
| [VertexAI Evaluation (gemini-1.5-pro-001)](/models/native-judgebench-vertexai-evaluation-gemini-1-5-pro-001) | Prompted Judge | 45.5 | 44.9 | 53.6 | 28.6 | 44.6 |
| [Arena-Hard (Llama-3.1-8B-Instruct)](/models/native-judgebench-arena-hard-llama-3-1-8b-instruct) | Prompted Judge | 38.3 | 45.9 | 44.6 | 33.3 | 40.9 |
| [Prometheus2-8x7b](/models/native-judgebench-prometheus2-8x7b) | Fine-Tuned Judge | 41.6 | 39.8 | 50.0 | 23.8 | 40.3 |
| [Arena-Hard (gemini-1.5-flash-001)](/models/native-judgebench-arena-hard-gemini-1-5-flash-001) | Prompted Judge | 42.9 | 36.7 | 50.0 | 21.4 | 39.7 |
| [Prometheus2-bgb-8x7b](/models/native-judgebench-prometheus2-bgb-8x7b) | Fine-Tuned Judge | 45.5 | 30.6 | 46.4 | 28.6 | 39.4 |
| [Auto-J](/models/native-judgebench-auto-j) | Fine-Tuned Judge | 40.3 | 29.6 | 44.6 | 28.6 | 36.6 |
| [JudgeLM-33B-v1.0](/models/native-judgebench-judgelm-33b-v1-0) | Fine-Tuned Judge | 32.5 | 49.0 | 33.9 | 19.0 | 35.7 |
| [Prometheus2-7b](/models/native-judgebench-prometheus2-7b) | Fine-Tuned Judge | 38.3 | 25.5 | 35.7 | 42.9 | 34.9 |
| [ChatEval (gpt-4o-2024-05-13)](/models/native-judgebench-chateval-gpt-4o-2024-05-13) | Multi-Agent Judge | 32.5 | 31.6 | 44.6 | 31.0 | 34.0 |
| [Arena-Hard (claude-3-haiku-20240307)](/models/native-judgebench-arena-hard-claude-3-haiku-20240307) | Prompted Judge | 35.1 | 34.7 | 33.9 | 21.4 | 33.1 |
| [JudgeLM-13B-v1.0](/models/native-judgebench-judgelm-13b-v1-0) | Fine-Tuned Judge | 26.6 | 29.6 | 28.6 | 19.0 | 26.9 |
| [JudgeLM-7B-v1.0](/models/native-judgebench-judgelm-7b-v1-0) | Fine-Tuned Judge | 23.4 | 29.6 | 32.1 | 11.9 | 25.1 |
| [PandaLM-7B-v1](/models/native-judgebench-pandalm-7b-v1) | Fine-Tuned Judge | 9.1 | 21.4 | 7.1 | 16.7 | 13.1 |

### Maintainer results on Claude response pairs

Claude-3-5-sonnet-20240620 response split. Nemotron supplement rows are omitted here because their response split is unspecified.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://huggingface.co/spaces/ScalerLab/JudgeBench)

| Judge | Type | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- | --- |
| [Arena-Hard (gpt-4o-2024-05-13)](/models/native-judgebench-arena-hard-gpt-4o-2024-05-13) | Prompted Judge | 52.6 | 51.0 | 61.8 | 38.7 | 51.9 |
| [Arena-Hard (Llama-3.1-70B-Instruct)](/models/native-judgebench-arena-hard-llama-3-1-70b-instruct) | Prompted Judge | 50.6 | 43.1 | 41.2 | 25.8 | 45.2 |
| [Arena-Hard (claude-3-5-sonnet-20240620)](/models/native-judgebench-arena-hard-claude-3-5-sonnet-20240620) | Prompted Judge | 42.2 | 52.9 | 55.9 | 32.3 | 44.8 |
| [Arena-Hard (Llama-3.1-8B-Instruct)](/models/native-judgebench-arena-hard-llama-3-1-8b-instruct) | Prompted Judge | 33.1 | 43.1 | 50.0 | 29.0 | 36.7 |
| [Arena-Hard (claude-3-haiku-20240307)](/models/native-judgebench-arena-hard-claude-3-haiku-20240307) | Prompted Judge | 37.7 | 29.4 | 32.4 | 9.7 | 32.2 |

### Maintainer Nemotron supplement — response split unspecified

The maintainer app inserts this CSV into both response-model views without a split field. These are 11 source-published results, not 22 independently measured rows.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://huggingface.co/spaces/ScalerLab/JudgeBench/blob/main/nemotron_results.csv)

| Judge | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| [Llama-3.3-Nemotron-Super-49B-GenRM](/models/native-judgebench-llama-3-3-nemotron-super-49b-genrm) | 71.4 | 73.5 | 87.5 | 76.2 | 75.1 |
| [Llama-3.3-Nemotron-Super-49B-GenRM + voting@32](/models/native-judgebench-llama-3-3-nemotron-super-49b-genrm-voting-32) | 70.8 | 83.7 | 87.5 | 83.3 | 78.6 |
| [Llama-3.3-Nemotron-Super-49B-GenRM-Multilingual](/models/native-judgebench-llama-3-3-nemotron-super-49b-genrm-multilingual) | 64.9 | 74.5 | 87.5 | 73.8 | 72.3 |
| [Llama-3.3-Nemotron-Super-49B-GenRM-Multilingual + voting@32](/models/native-judgebench-llama-3-3-nemotron-super-49b-genrm-multilingual-voting-32) | 65.6 | 82.7 | 87.5 | 85.7 | 76.3 |
| [Llama-3.3-Nemotron-70B-Reward](/models/native-judgebench-llama-3-3-nemotron-70b-reward) | 70.8 | 76.5 | 82.1 | 66.7 | 73.7 |
| [Llama-3.3-Nemotron-70B-Reward-Multilingual](/models/native-judgebench-llama-3-3-nemotron-70b-reward-multilingual) | 66.2 | 71.4 | 82.1 | 59.5 | 69.4 |
| [Llama-3.1-Nemotron-70B-Reward](/models/native-judgebench-llama-3-1-nemotron-70b-reward) | 62.3 | 72.5 | 76.8 | 57.1 | 66.9 |
| [Qwen-3-Nemotron-32B-Reward](/models/native-judgebench-qwen-3-nemotron-32b-reward) | 70.1 | 67.4 | 78.6 | 83.3 | 72.3 |
| [Qwen-2.5-Nemotron-32B-Reward](/models/native-judgebench-qwen-2-5-nemotron-32b-reward) | 61.7 | 74.5 | 76.2 | 82.1 | 70.3 |
| [Qwen3-Nemotron-32B-GenRM-Principle](/models/native-judgebench-qwen3-nemotron-32b-genrm-principle) | 74.6 | 85.7 | 85.7 | 90.5 | 81.4 |
| [Llama-3.3-Nemotron-70B-Reward-Principle](/models/native-judgebench-llama-3-3-nemotron-70b-reward-principle) | 74.0 | 74.5 | 82.1 | 81.0 | 76.3 |

### Paper: judge methods

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2)

| Judge | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| Prompted Judges |  |  |  |  |  |
| [Vanilla (GPT-4o)](/models/native-judgebench-vanilla-gpt-4o-2024-05-13) | 44.16 | 47.96 | 66.07 | 61.90 | 50.86 |
| [Arena-Hard Judge (GPT-4o)](/models/native-judgebench-arena-hard-gpt-4o-2024-05-13) | 50.65 | 54.08 | 75.00 | 59.52 | 56.57 |
| [VertexAI Evaluation (Gemini-1.5-pro)](/models/native-judgebench-vertexai-evaluation-gemini-1-5-pro-001) | 45.45 | 44.90 | 53.57 | 28.57 | 44.57 |
| Fine-tuned Judges |  |  |  |  |  |
| [PandaLM](/models/native-judgebench-pandalm-7b-v1) | 9.09 | 21.43 | 7.14 | 16.67 | 13.14 |
| [Prometheus2-7b](/models/native-judgebench-prometheus2-7b) | 38.31 | 25.51 | 35.71 | 42.86 | 34.86 |
| [Prometheus2-8x7b](/models/native-judgebench-prometheus2-8x7b) | 41.56 | 39.80 | 50.00 | 23.81 | 40.29 |
| [Prometheus2-bgb-8x7b](/models/native-judgebench-prometheus2-bgb-8x7b) | 45.45 | 30.61 | 46.43 | 28.57 | 39.43 |
| [JudgeLM-7B](/models/native-judgebench-judgelm-7b-v1-0) | 23.38 | 29.59 | 32.14 | 11.90 | 25.14 |
| [JudgeLM-13B](/models/native-judgebench-judgelm-13b-v1-0) | 26.62 | 29.59 | 28.57 | 19.05 | 26.86 |
| [JudgeLM-33B](/models/native-judgebench-judgelm-33b-v1-0) | 32.47 | 48.98 | 33.93 | 19.05 | 35.71 |
| [AutoJ](/models/native-judgebench-auto-j) | 40.26 | 29.59 | 44.64 | 28.57 | 36.57 |
| [Skywork-LLaMA-3.1B-8B](/models/native-judgebench-skywork-critic-llama-3-1-8b) | 51.30 | 54.08 | 73.21 | 33.33 | 53.43 |
| [Skywork-LLaMA-3.1B-70B](/models/native-judgebench-skywork-critic-llama-3-1-70b) | 55.84 | 55.10 | 73.21 | 47.62 | 57.43 |
| Multi-Agent Judges |  |  |  |  |  |
| [ChatEval](/models/native-judgebench-chateval-gpt-4o-2024-05-13) | 32.47 | 31.63 | 44.64 | 30.95 | 34.00 |

### Paper: Arena-Hard prompted judges

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2)

| Judge model | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| [GPT-4o](/models/native-judgebench-arena-hard-gpt-4o-2024-05-13) | 50.65 | 54.08 | 75.00 | 59.52 | 56.57 |
| [GPT-4o-mini](/models/native-judgebench-arena-hard-gpt-4o-mini-2024-07-18) | 48.05 | 43.88 | 69.64 | 45.24 | 50.00 |
| [o1-preview](/models/native-judgebench-arena-hard-o1-preview-2024-09-12) | 66.23 | 79.59 | 85.71 | 85.71 | 75.43 |
| [o1-mini](/models/native-judgebench-arena-hard-o1-mini-2024-09-12) | 58.44 | 62.24 | 82.14 | 78.57 | 65.71 |
| [o3-mini (high)](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-high) | 67.53 | 89.80 | 87.50 | 100.0 | 80.86 |
| [o3-mini (medium)](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-medium) | 62.34 | 86.73 | 85.71 | 92.86 | 76.57 |
| [o3-mini (low)](/models/native-judgebench-arena-hard-o3-mini-2025-01-31-low) | 62.99 | 69.39 | 83.93 | 83.33 | 70.57 |
| [Claude-3.5-Sonnet](/models/native-judgebench-arena-hard-claude-3-5-sonnet-20240620) | 62.34 | 66.33 | 66.07 | 64.29 | 64.29 |
| [Claude-3-Haiku](/models/native-judgebench-arena-hard-claude-3-haiku-20240307) | 35.06 | 34.69 | 33.93 | 21.43 | 33.14 |
| [Llama-3.1-405B-Instruct](/models/native-judgebench-arena-hard-llama-3-1-405b-instruct) | 55.84 | 54.08 | 69.64 | 50.00 | 56.86 |
| [Llama-3.1-70B-Instruct](/models/native-judgebench-arena-hard-llama-3-1-70b-instruct) | 51.30 | 48.98 | 60.71 | 52.38 | 52.29 |
| [Llama-3.1-8B-Instruct](/models/native-judgebench-arena-hard-llama-3-1-8b-instruct) | 38.31 | 45.92 | 44.64 | 33.33 | 40.86 |
| [Gemini-1.5-pro](/models/native-judgebench-arena-hard-gemini-1-5-pro-001) | 49.35 | 42.86 | 64.29 | 26.19 | 47.14 |
| [Gemini-1.5-flash](/models/native-judgebench-arena-hard-gemini-1-5-flash-001) | 42.86 | 36.73 | 50.00 | 21.43 | 39.71 |
| [Deepseek-R1](/models/native-judgebench-arena-hard-deepseek-r1-250120) | 59.09 | 82.65 | 80.36 | 92.86 | 73.14 |

### Paper: reward models

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2)

| Reward model | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| [Skywork-Reward-Gemma-2-27B](/models/native-judgebench-skywork-reward-gemma-2-27b) | 59.74 | 66.33 | 83.93 | 50.00 | 64.29 |
| [Skywork-Reward-Llama-3.1-8B](/models/native-judgebench-skywork-reward-llama-3-1-8b) | 59.09 | 64.29 | 76.79 | 50.00 | 62.29 |
| [InternLM2-20B-Reward](/models/native-judgebench-internlm2-20b-reward) | 62.34 | 69.39 | 66.07 | 50.00 | 63.43 |
| [InternLM2-7B-Reward](/models/native-judgebench-internlm2-7b-reward) | 56.49 | 61.22 | 71.43 | 50.00 | 59.43 |
| [GRM-Gemma-2B](/models/native-judgebench-grm-gemma-2b) | 62.99 | 53.06 | 64.29 | 54.76 | 59.43 |

### Paper: solving versus judging

Paper v2; configurations and precision remain as published. Solver results are direct answers rather than judge preferences.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2)

| Setup | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| [GPT-4o Solver](/models/native-judgebench-gpt-4o-solver) | 48.70 | 53.06 | 58.93 | 73.81 | 54.57 |
| [GPT-4o Judge](/models/native-judgebench-arena-hard-gpt-4o-2024-05-13) | 50.65 | 54.08 | 75.00 | 59.52 | 56.57 |
| [Claude-3.5-Sonnet Solver](/models/native-judgebench-claude-3-5-sonnet-solver) | 61.04 | 62.24 | 60.71 | 88.10 | 64.57 |
| [Claude-3.5-Sonnet Judge](/models/native-judgebench-arena-hard-claude-3-5-sonnet-20240620) | 62.34 | 66.33 | 66.07 | 64.29 | 64.29 |
| [Llama-3.1-405B-Instruct Solver](/models/native-judgebench-llama-3-1-405b-instruct-solver) | 48.05 | 67.86 | 63.27 | 66.67 | 57.71 |
| [Llama-3.1-405B-Instruct Judge](/models/native-judgebench-arena-hard-llama-3-1-405b-instruct) | 55.84 | 54.08 | 69.64 | 50.00 | 56.86 |
| [Gemini-1.5-pro Solver](/models/native-judgebench-gemini-1-5-pro-solver) | 33.12 | 42.86 | 37.50 | 64.29 | 40.29 |
| [Gemini-1.5-pro Judge](/models/native-judgebench-arena-hard-gemini-1-5-pro-001) | 49.35 | 42.86 | 64.29 | 26.19 | 47.14 |

### Paper: judgment outcomes

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2)

| Judge | A>B count | A<B count | Tie count | Invalid count |
| --- | --- | --- | --- | --- |
| [PandaLM-7B](/models/native-judgebench-pandalm-7b-v1) | 45 | 114 | 479 | 62 |
| [Prometheus2-7b](/models/native-judgebench-prometheus2-7b) | 395 | 232 | 0 | 73 |
| [Prometheus2-8x7b](/models/native-judgebench-prometheus2-8x7b) | 331 | 328 | 0 | 41 |
| [Prometheus2-bgb-8x7b](/models/native-judgebench-prometheus2-bgb-8x7b) | 239 | 215 | 0 | 246 |
| [JudgeLM-7B](/models/native-judgebench-judgelm-7b-v1-0) | 399 | 229 | 72 | 0 |
| [JudgeLM-13B](/models/native-judgebench-judgelm-13b-v1-0) | 355 | 312 | 33 | 0 |
| [JudgeLM-33B](/models/native-judgebench-judgelm-33b-v1-0) | 344 | 264 | 92 | 0 |
| [AutoJ](/models/native-judgebench-auto-j) | 289 | 378 | 33 | 0 |
| [Skywork-LLaMA-3.1B-8B](/models/native-judgebench-skywork-critic-llama-3-1-8b) | 346 | 354 | 0 | 0 |
| [Skywork-LLaMA-3.1B-70B](/models/native-judgebench-skywork-critic-llama-3-1-70b) | 390 | 310 | 0 | 0 |

### Paper: order inconsistency

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2)

| Judge | Inconsistent (%) |
| --- | --- |
| [PandaLM-7B](/models/native-judgebench-pandalm-7b-v1) | 29.14% |
| [Prometheus2-7b](/models/native-judgebench-prometheus2-7b) | 52.29% |
| [Prometheus2-8x7b](/models/native-judgebench-prometheus2-8x7b) | 40.00% |
| [Prometheus2-bgb-8x7b](/models/native-judgebench-prometheus2-bgb-8x7b) | 43.71% |
| [JudgeLM-7B](/models/native-judgebench-judgelm-7b-v1-0) | 59.71% |
| [JudgeLM-13B](/models/native-judgebench-judgelm-13b-v1-0) | 54.57% |
| [JudgeLM-33B](/models/native-judgebench-judgelm-33b-v1-0) | 38.00% |
| [AutoJ](/models/native-judgebench-auto-j) | 43.71% |
| [Skywork-Llama-3.1B-8B](/models/native-judgebench-skywork-critic-llama-3-1-8b) | 18.86% |
| [Skywork-Llama-3.1B-70B](/models/native-judgebench-skywork-critic-llama-3-1-70b) | 18.29% |

### Paper: Prometheus and its base model

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2)

| Judge | Knowledge accuracy (%) |
| --- | --- |
| [Prometheus2-7b](/models/native-judgebench-prometheus2-7b) | 38.31 |
| [Vanilla (Mistral-7B-v0.1-Instruct)](/models/native-judgebench-vanilla-mistral-7b-v0-1-instruct) | 7.43 |
| [Arena-Hard (Mistral-7B-v0.1-Instruct)](/models/native-judgebench-arena-hard-mistral-7b-v0-1-instruct) | 6.57 |

### Paper: original versus augmented knowledge

Paper v2; source group labels and published precision are retained.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/2410.12784v2)

| Judge model | Original accuracy (%) and rank | Augmented accuracy (%) and rank |
| --- | --- | --- |
| [gpt-4o](/models/native-judgebench-arena-hard-gpt-4o-2024-05-13) | 50.65 (3rd) | 46.49 (3rd) |
| [gpt-4o-mini](/models/native-judgebench-arena-hard-gpt-4o-mini-2024-07-18) | 48.05 (4th) | 44.03 (4th) |
| [claude-3.5-sonnet](/models/native-judgebench-arena-hard-claude-3-5-sonnet-20240620) | 62.34 (1st) | 63.25 (1st) |
| [claude-3-haiku](/models/native-judgebench-arena-hard-claude-3-haiku-20240307) | 35.06 (6th) | 39.35 (6th) |
| [llama-3.1-70b-instruct](/models/native-judgebench-arena-hard-llama-3-1-70b-instruct) | 51.30 (2nd) | 52.60 (2nd) |
| [llama-3.1-8b-instruct](/models/native-judgebench-arena-hard-llama-3-1-8b-instruct) | 38.31 (5th) | 40.00 (5th) |

## FAQ

### Which JudgeBench results are included?

This page includes 11 numeric result and configuration tables from JudgeBench’s original paper and available benchmark-owner updates, covering 54 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

### Can I compare these JudgeBench scores with the Perplexity panel?

Compare JudgeBench scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

### Where can I download the JudgeBench results?

The results download on this page provides every imported JudgeBench table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
