Skip to main content
BenchLM

JudgeBench

We show this table for reference; we do not rank on it.

JudgeBench evaluates whether a judge can identify the better response in pairs with objective correctness labels. It tests judgment on challenging questions rather than agreement with writing style preferences.

Original benchmark results

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

11 source tables · 54 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)

Maintainer results on GPT-4o response pairs

GPT-4o-2024-05-13 response split. The public app rounds to one decimal; paper tables retain two decimals.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Maintainer results on GPT-4o response pairs. Published source metrics with original precision; unreported cells are not zero.
JudgeTypeKnowledge (%)Reasoning (%)Math (%)Coding (%)Overall (%)
Arena-Hard (o3-mini-2025-01-31 (high))Prompted Judge67.589.887.5100.080.9
Arena-Hard (o3-mini-2025-01-31 (medium))Prompted Judge62.386.785.792.976.6
Arena-Hard (o1-preview-2024-09-12)Prompted Judge66.279.685.785.775.4
Arena-Hard (DeepSeek-R1-250120)Prompted Judge59.182.780.492.973.1
Arena-Hard (o3-mini-2025-01-31 (low))Prompted Judge63.069.483.983.370.6
Arena-Hard (o1-mini-2024-09-12)Prompted Judge58.462.282.178.665.7
Arena-Hard (claude-3-5-sonnet-20240620)Prompted Judge62.366.366.164.364.3
Skywork-Reward-Gemma-2-27BReward Model59.766.383.950.064.3
InternLM2-20B-RewardReward Model62.369.466.150.063.4
CompassJudger-1-32BFine-Tuned Judge58.462.275.059.562.3
Skywork-Reward-Llama-3.1-8BReward Model59.164.376.850.062.3
GRM-Gemma-2BReward Model63.053.164.354.859.4
InternLM2-7B-RewardReward Model56.561.271.450.059.4
Skywork-Critic-Llama-3.1-70BFine-Tuned Judge55.855.173.247.657.4
Arena-Hard (Llama-3.1-405B-Instruct)Prompted Judge55.854.169.650.056.9
Arena-Hard (gpt-4o-2024-05-13)Prompted Judge50.654.175.059.556.6
CompassJudger-1-14BFine-Tuned Judge51.960.267.942.955.7
Skywork-Critic-Llama-3.1-8BFine-Tuned Judge51.354.173.233.353.4
Arena-Hard (Llama-3.1-70B-Instruct)Prompted Judge51.349.060.752.452.3
Vanilla (gpt-4o-2024-05-13)Prompted Judge44.248.066.161.950.9
Arena-Hard (gpt-4o-mini-2024-07-18)Prompted Judge48.143.969.645.250.0
Arena-Hard (gemini-1.5-pro-001)Prompted Judge49.442.964.326.247.1
CompassJudger-1-7BFine-Tuned Judge42.237.869.647.646.0
CompassJudger-1-1.5BFine-Tuned Judge48.139.850.035.744.6
VertexAI Evaluation (gemini-1.5-pro-001)Prompted Judge45.544.953.628.644.6
Arena-Hard (Llama-3.1-8B-Instruct)Prompted Judge38.345.944.633.340.9
Prometheus2-8x7bFine-Tuned Judge41.639.850.023.840.3
Arena-Hard (gemini-1.5-flash-001)Prompted Judge42.936.750.021.439.7
Prometheus2-bgb-8x7bFine-Tuned Judge45.530.646.428.639.4
Auto-JFine-Tuned Judge40.329.644.628.636.6
JudgeLM-33B-v1.0Fine-Tuned Judge32.549.033.919.035.7
Prometheus2-7bFine-Tuned Judge38.325.535.742.934.9
ChatEval (gpt-4o-2024-05-13)Multi-Agent Judge32.531.644.631.034.0
Arena-Hard (claude-3-haiku-20240307)Prompted Judge35.134.733.921.433.1
JudgeLM-13B-v1.0Fine-Tuned Judge26.629.628.619.026.9
JudgeLM-7B-v1.0Fine-Tuned Judge23.429.632.111.925.1
PandaLM-7B-v1Fine-Tuned Judge9.121.47.116.713.1

Maintainer results on Claude response pairs

Claude-3-5-sonnet-20240620 response split. Nemotron supplement rows are omitted here because their response split is unspecified.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Maintainer results on Claude response pairs. Published source metrics with original precision; unreported cells are not zero.
JudgeTypeKnowledge (%)Reasoning (%)Math (%)Coding (%)Overall (%)
Arena-Hard (gpt-4o-2024-05-13)Prompted Judge52.651.061.838.751.9
Arena-Hard (Llama-3.1-70B-Instruct)Prompted Judge50.643.141.225.845.2
Arena-Hard (claude-3-5-sonnet-20240620)Prompted Judge42.252.955.932.344.8
Arena-Hard (Llama-3.1-8B-Instruct)Prompted Judge33.143.150.029.036.7
Arena-Hard (claude-3-haiku-20240307)Prompted Judge37.729.432.49.732.2

Maintainer Nemotron supplement — response split unspecified

The maintainer app inserts this CSV into both response-model views without a split field. These are 11 source-published results, not 22 independently measured rows.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Maintainer Nemotron supplement — response split unspecified. Published source metrics with original precision; unreported cells are not zero.
JudgeKnowledge (%)Reasoning (%)Math (%)Coding (%)Overall (%)
Llama-3.3-Nemotron-Super-49B-GenRM71.473.587.576.275.1
Llama-3.3-Nemotron-Super-49B-GenRM + voting@3270.883.787.583.378.6
Llama-3.3-Nemotron-Super-49B-GenRM-Multilingual64.974.587.573.872.3
Llama-3.3-Nemotron-Super-49B-GenRM-Multilingual + voting@3265.682.787.585.776.3
Llama-3.3-Nemotron-70B-Reward70.876.582.166.773.7
Llama-3.3-Nemotron-70B-Reward-Multilingual66.271.482.159.569.4
Llama-3.1-Nemotron-70B-Reward62.372.576.857.166.9
Qwen-3-Nemotron-32B-Reward70.167.478.683.372.3
Qwen-2.5-Nemotron-32B-Reward61.774.576.282.170.3
Qwen3-Nemotron-32B-GenRM-Principle74.685.785.790.581.4
Llama-3.3-Nemotron-70B-Reward-Principle74.074.582.181.076.3

Paper: judge methods

Paper v2; source group labels and published precision are retained.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper: judge methods. Published source metrics with original precision; unreported cells are not zero.
JudgeKnowledge (%)Reasoning (%)Math (%)Coding (%)Overall (%)
Prompted Judges
Vanilla (GPT-4o)44.1647.9666.0761.9050.86
Arena-Hard Judge (GPT-4o)50.6554.0875.0059.5256.57
VertexAI Evaluation (Gemini-1.5-pro)45.4544.9053.5728.5744.57
Fine-tuned Judges
PandaLM9.0921.437.1416.6713.14
Prometheus2-7b38.3125.5135.7142.8634.86
Prometheus2-8x7b41.5639.8050.0023.8140.29
Prometheus2-bgb-8x7b45.4530.6146.4328.5739.43
JudgeLM-7B23.3829.5932.1411.9025.14
JudgeLM-13B26.6229.5928.5719.0526.86
JudgeLM-33B32.4748.9833.9319.0535.71
AutoJ40.2629.5944.6428.5736.57
Skywork-LLaMA-3.1B-8B51.3054.0873.2133.3353.43
Skywork-LLaMA-3.1B-70B55.8455.1073.2147.6257.43
Multi-Agent Judges
ChatEval32.4731.6344.6430.9534.00

Paper: Arena-Hard prompted judges

Paper v2; source group labels and published precision are retained.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper: Arena-Hard prompted judges. Published source metrics with original precision; unreported cells are not zero.
Judge modelKnowledge (%)Reasoning (%)Math (%)Coding (%)Overall (%)
GPT-4o50.6554.0875.0059.5256.57
GPT-4o-mini48.0543.8869.6445.2450.00
o1-preview66.2379.5985.7185.7175.43
o1-mini58.4462.2482.1478.5765.71
o3-mini (high)67.5389.8087.50100.080.86
o3-mini (medium)62.3486.7385.7192.8676.57
o3-mini (low)62.9969.3983.9383.3370.57
Claude-3.5-Sonnet62.3466.3366.0764.2964.29
Claude-3-Haiku35.0634.6933.9321.4333.14
Llama-3.1-405B-Instruct55.8454.0869.6450.0056.86
Llama-3.1-70B-Instruct51.3048.9860.7152.3852.29
Llama-3.1-8B-Instruct38.3145.9244.6433.3340.86
Gemini-1.5-pro49.3542.8664.2926.1947.14
Gemini-1.5-flash42.8636.7350.0021.4339.71
Deepseek-R159.0982.6580.3692.8673.14

Paper: reward models

Paper v2; source group labels and published precision are retained.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper: reward models. Published source metrics with original precision; unreported cells are not zero.
Reward modelKnowledge (%)Reasoning (%)Math (%)Coding (%)Overall (%)
Skywork-Reward-Gemma-2-27B59.7466.3383.9350.0064.29
Skywork-Reward-Llama-3.1-8B59.0964.2976.7950.0062.29
InternLM2-20B-Reward62.3469.3966.0750.0063.43
InternLM2-7B-Reward56.4961.2271.4350.0059.43
GRM-Gemma-2B62.9953.0664.2954.7659.43

Paper: solving versus judging

Paper v2; configurations and precision remain as published. Solver results are direct answers rather than judge preferences.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper: solving versus judging. Published source metrics with original precision; unreported cells are not zero.
SetupKnowledge (%)Reasoning (%)Math (%)Coding (%)Overall (%)
GPT-4o Solver48.7053.0658.9373.8154.57
GPT-4o Judge50.6554.0875.0059.5256.57
Claude-3.5-Sonnet Solver61.0462.2460.7188.1064.57
Claude-3.5-Sonnet Judge62.3466.3366.0764.2964.29
Llama-3.1-405B-Instruct Solver48.0567.8663.2766.6757.71
Llama-3.1-405B-Instruct Judge55.8454.0869.6450.0056.86
Gemini-1.5-pro Solver33.1242.8637.5064.2940.29
Gemini-1.5-pro Judge49.3542.8664.2926.1947.14

Paper: judgment outcomes

Paper v2; source group labels and published precision are retained.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper: judgment outcomes. Published source metrics with original precision; unreported cells are not zero.
JudgeA>B countA<B countTie countInvalid count
PandaLM-7B4511447962
Prometheus2-7b395232073
Prometheus2-8x7b331328041
Prometheus2-bgb-8x7b2392150246
JudgeLM-7B399229720
JudgeLM-13B355312330
JudgeLM-33B344264920
AutoJ289378330
Skywork-LLaMA-3.1B-8B34635400
Skywork-LLaMA-3.1B-70B39031000

Paper: order inconsistency

Paper v2; source group labels and published precision are retained.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper: order inconsistency. Published source metrics with original precision; unreported cells are not zero.
JudgeInconsistent (%)
PandaLM-7B29.14%
Prometheus2-7b52.29%
Prometheus2-8x7b40.00%
Prometheus2-bgb-8x7b43.71%
JudgeLM-7B59.71%
JudgeLM-13B54.57%
JudgeLM-33B38.00%
AutoJ43.71%
Skywork-Llama-3.1B-8B18.86%
Skywork-Llama-3.1B-70B18.29%

Paper: Prometheus and its base model

Paper v2; source group labels and published precision are retained.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper: Prometheus and its base model. Published source metrics with original precision; unreported cells are not zero.
JudgeKnowledge accuracy (%)
Prometheus2-7b38.31
Vanilla (Mistral-7B-v0.1-Instruct)7.43
Arena-Hard (Mistral-7B-v0.1-Instruct)6.57

Paper: original versus augmented knowledge

Paper v2; source group labels and published precision are retained.

Published source

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Paper: original versus augmented knowledge. Published source metrics with original precision; unreported cells are not zero.
Judge modelOriginal accuracy (%) and rankAugmented accuracy (%) and rank
gpt-4o50.65 (3rd)46.49 (3rd)
gpt-4o-mini48.05 (4th)44.03 (4th)
claude-3.5-sonnet62.34 (1st)63.25 (1st)
claude-3-haiku35.06 (6th)39.35 (6th)
llama-3.1-70b-instruct51.30 (2nd)52.60 (2nd)
llama-3.1-8b-instruct38.31 (5th)40.00 (5th)

Original benchmark results and configurations

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.

Snapshot

11 source tables54 configurationsDisplay only

About JudgeBench

Year

2024

Tasks

Choosing the better of two model responses

Format

Decision classification

Difficulty

Depends on the evaluated split and label mapping

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

Freshness and provenance

Version

JudgeBench 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Which JudgeBench results are included?

This page includes 11 numeric result and configuration tables from JudgeBench’s original paper and available benchmark-owner updates, covering 54 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

Can I compare these JudgeBench scores with the Perplexity panel?

Compare JudgeBench scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

Where can I download the JudgeBench results?

The results download on this page provides every imported JudgeBench table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.

Last updated: October 1, 2026 source review · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.