# Llama-3.3-Nemotron-Super-49B-GenRM + voting@32 (JudgeBench) Benchmark Scores & Performance

> Llama-3.3-Nemotron-Super-49B-GenRM + voting@32 (JudgeBench) has published results in 1 original benchmark tables. The evaluated configuration, source metrics, and precision remain visible below. These results do not produce a general model score or rank.

Canonical page: https://benchlm.ai/models/native-judgebench-llama-3-3-nemotron-super-49b-genrm-voting-32

Last updated: 2026-10-01

General benchmark catalog last updated: October 1, 2026. This profile’s source review has its own date above.

## Model Details

| Property | Value |
|----------|-------|
| Creator | NVIDIA |
| Source Type | Evaluated configuration |
| Reasoning Type | Unspecified |
| Context Window | Not established by evaluation |
| Official model card | [Benchmark-owner results and configuration](https://huggingface.co/spaces/ScalerLab/JudgeBench) |
| Overall Score | Not computed (source protocol results only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Llama-3.3-Nemotron-Super-49B-GenRM + voting@32 (JudgeBench)
- Variant: benchmark-system
- Benchmarks covered: 0 of 645
- Coverage note: Original benchmark result tables appear below; these metrics are separate from weighted benchmark slots.

## Original benchmark results

[All model results (JSON)](/api/data/benchmarks?model=native-judgebench-llama-3-3-nemotron-super-49b-genrm-voting-32) · [Numeric metrics (CSV)](/api/data/benchmarks?model=native-judgebench-llama-3-3-nemotron-super-49b-genrm-voting-32&format=csv)

### Maintainer Nemotron supplement — response split unspecified

JudgeBench’s paper and maintained leaderboard report judges under different prompts, reasoning efforts, and response-model splits. Accuracy is a percentage. The paper’s 700-pair GPT-4o split is distinct from the 270-pair Claude split. The app appends the same Nemotron CSV to both tabs without specifying a response split; we show that supplement once with its split unresolved. Live one-decimal values and paper two-decimal values are preserved separately.

The maintainer app inserts this CSV into both response-model views without a split field. These are 11 source-published results, not 22 independently measured rows.

JudgeBench: A Benchmark for Evaluating LLM-Based Judges — Sijun Tan and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://huggingface.co/spaces/ScalerLab/JudgeBench/blob/main/nemotron_results.csv) · [Full JudgeBench results](/benchmarks/judgebench)

| Judge | Knowledge (%) | Reasoning (%) | Math (%) | Coding (%) | Overall (%) |
| --- | --- | --- | --- | --- | --- |
| Llama-3.3-Nemotron-Super-49B-GenRM + voting@32 | 70.8 | 83.7 | 87.5 | 83.3 | 78.6 |

## Other NVIDIA Models

- [Nemotron 3 Ultra](/models/nemotron-3-ultra) - Score: 45.35
- [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) - Score: 31.28
- [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) - Score: 20.13
- [Nemotron-4 15B](/models/nemotron-4-15b) - Score: 10.05
- [Nemotron 3 Nano 30B](/models/nemotron-3-nano-30b) - Score: not computed
- [Nemotron 3 Super 120B A12B](/models/nemotron-3-super-120b-a12b) - Score: not computed
- [Nemotron 3 Super 100B](/models/nemotron-3-super-100b) - Score: not computed
- [Nemotron Ultra 253B](/models/nemotron-ultra-253b) - Score: not computed
- [Cosmos3-Edge](/models/cosmos3-edge) - Score: not computed
- [Audio Flamingo 3 7B](/models/audio-flamingo-3-7b) - Score: not computed
- [Canary 180M Flash](/models/canary-180m-flash) - Score: not computed
- [Llama-3.1-Nemotron-70B-Reward (JudgeBench)](/models/native-judgebench-llama-3-1-nemotron-70b-reward) - Score: not computed
- [Llama-3.3-Nemotron-70B-Reward (JudgeBench)](/models/native-judgebench-llama-3-3-nemotron-70b-reward) - Score: not computed
- [Llama-3.3-Nemotron-70B-Reward-Multilingual (JudgeBench)](/models/native-judgebench-llama-3-3-nemotron-70b-reward-multilingual) - Score: not computed
- [Llama-3.3-Nemotron-70B-Reward-Principle (JudgeBench)](/models/native-judgebench-llama-3-3-nemotron-70b-reward-principle) - Score: not computed
- [Llama-3.3-Nemotron-Super-49B-GenRM (JudgeBench)](/models/native-judgebench-llama-3-3-nemotron-super-49b-genrm) - Score: not computed
- [Llama-3.3-Nemotron-Super-49B-GenRM-Multilingual (JudgeBench)](/models/native-judgebench-llama-3-3-nemotron-super-49b-genrm-multilingual) - Score: not computed
- [Llama-3.3-Nemotron-Super-49B-GenRM-Multilingual + voting@32 (JudgeBench)](/models/native-judgebench-llama-3-3-nemotron-super-49b-genrm-multilingual-voting-32) - Score: not computed
- [Nemotron 3.5 ASR Streaming 0.6B](/models/nemotron-3-5-asr-streaming-0-6b) - Score: not computed
- [Parakeet CTC 1.1B](/models/parakeet-ctc-1-1b) - Score: not computed
- [Parakeet TDT 0.6B v3](/models/parakeet-tdt-0-6b-v3) - Score: not computed
- [Qwen-2.5-Nemotron-32B-Reward (JudgeBench)](/models/native-judgebench-qwen-2-5-nemotron-32b-reward) - Score: not computed
- [Qwen-3-Nemotron-32B-Reward (JudgeBench)](/models/native-judgebench-qwen-3-nemotron-32b-reward) - Score: not computed
- [Qwen3-Nemotron-32B-GenRM-Principle (JudgeBench)](/models/native-judgebench-qwen3-nemotron-32b-genrm-principle) - Score: not computed
