# Grok 4.20 Benchmark Scores & Performance

> Grok 4.20 by xAI scores 67.13/100 overall, ranking #29 out of 483 AI models.

Canonical page: https://benchlm.ai/models/grok-4-20-beta

Last updated: September 10, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | xAI |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 2M |
| Overall Score | 67.13/100 |
| Overall Rank | #29 of 483 |

## Family & Coverage

- Family: Grok 4.20
- Variant: reasoning
- Benchmarks covered: 24 of 435
- Sibling models: [Grok 4.20 Multi-agent](/models/grok-4-20-multi-agent-beta)
- Related earlier model: [Grok 4.1](/models/grok-4-1)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 47.1% |
| [DeepSearchQA](/benchmarks/deepsearchqa) | 62.8% |
| [Gert Labs](/benchmarks/gertlabs) | 38.36% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 44.2% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [LiveCodeBench Pro](/benchmarks/livecodebench-pro) | 74.2% |
| [SWE-bench Verified](/benchmarks/swe-bench-verified) | 76.7% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 51.8% |
| [Vibe Code Bench](/benchmarks/vibecodebench) | 4.06% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 84.3% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 72.2% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMMU-Pro](/benchmarks/mmmu-pro) | 75.2% |
| [CharXiv](/benchmarks/charxiv) | 60.9% |
| [ERQA](/benchmarks/erqa) | 54.1% |
| [SimpleVQA](/benchmarks/simplevqa) | 57.4% |
| [MedXpertQA (MM)](/benchmarks/medxpertqamm) | 65.8% |
| [Design Arena Website](/benchmarks/designarenawebsite) | 1242 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [ARC-AGI-2](/benchmarks/arc-agi-2) | 53.3% |
| [ARC-AGI-3](/benchmarks/arcagi3) | 0.1% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA-D](/benchmarks/gpqa-diamond) | 88.5% |
| [HLE w/o tools](/benchmarks/hlenotools) | 31.6% |
| [HealthBench Hard](/benchmarks/healthbench-hard) | 20.3% |
| [MedXpertQA (Text)](/benchmarks/medxpertqatext) | 50.2% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 88.6% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 86.3% |

## Other xAI Models

- [Grok 4.6](/models/grok-4-6) - Score: 70.08
- [Grok 4.5](/models/grok-4-5) - Score: 68.1
- [Grok 4.3](/models/grok-4-3) - Score: 59.73
- [Grok 4.1](/models/grok-4-1) - Score: 57.6
- [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) - Score: 55
- [Grok 4](/models/grok-4) - Score: 53.38
- [Grok 3 [Beta]](/models/grok-3-beta) - Score: 51.71
- [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) - Score: 51.31
- [Grok 4.1 Fast](/models/grok-4-1-fast) - Score: 47.26
- [Grok 3 Mini](/models/grok-3-mini) - Score: 46.52
- [Grok Code Fast 1](/models/grok-code-fast-1) - Score: 34.16
- [Grok Build 0.1](/models/grok-build-0-1) - Score: not computed
- [Grok TTS](/models/grok-tts) - Score: not computed
- [Grok 4.20 Multi-agent](/models/grok-4-20-multi-agent-beta) - Score: not computed
- [Grok Realtime](/models/grok-realtime) - Score: not computed
- [Grok Voice Think Fast 1.0](/models/grok-voice-think-fast-1-0) - Score: not computed
- [Grok Voice Think Fast 2.0](/models/grok-voice-think-fast-2-0) - Score: not computed
