# Grok 4.3 Benchmark Scores & Performance

> Grok 4.3 by xAI scores 59.73/100 overall, ranking #61 out of 483 AI models.

Canonical page: https://benchlm.ai/models/grok-4-3

Last updated: September 10, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | xAI |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Overall Score | 59.73/100 |
| Overall Rank | #61 of 483 |

## Family & Coverage

- Family: Grok 4.3
- Variant: base
- Benchmarks covered: 30 of 435
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [τ²-bench results](/benchmarks/tau2-bench) | 97.7% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 29.2% |
| [APEX-Agents-AA](/benchmarks/apexagentsaa) | 17.0% |
| [Gert Labs](/benchmarks/gertlabs) | 43.86% |
| [ResearchClawBench](/benchmarks/researchclawbench) | 12.4% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 41.9% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 17.2% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1018 |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [SciCode](/benchmarks/scicode) | 47.3% |
| [AA-SciCode](/benchmarks/aascicode) | 48.3% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 84.5% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 71.4% |
| [AA Coding Index](/benchmarks/aacodingindex) | 42.3% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMMU-Pro](/benchmarks/mmmu-pro) | 78.1% |
| [Design Arena Website](/benchmarks/designarenawebsite) | 1206 |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 78.1% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 64.3% |
| [CritPt](/benchmarks/critpt) | 8.0% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 37.6% |
| [GPQA](/benchmarks/gpqa) | 90.1% |
| [HLE](/benchmarks/hle) | 35% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 34.6% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 25.0% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 90.1% |
| [AA-HLE](/benchmarks/aahle) | 37.2% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 18.0% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 91.4% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 85.8% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [IFBench](/benchmarks/ifbench) | 81.3% |
| [AA-IFBench](/benchmarks/aaifbench) | 81.3% |

## Other xAI Models

- [Grok 4.6](/models/grok-4-6) - Score: 70.08
- [Grok 4.5](/models/grok-4-5) - Score: 68.1
- [Grok 4.20](/models/grok-4-20-beta) - Score: 67.13
- [Grok 4.1](/models/grok-4-1) - Score: 57.6
- [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) - Score: 55
- [Grok 4](/models/grok-4) - Score: 53.38
- [Grok 3 [Beta]](/models/grok-3-beta) - Score: 51.71
- [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) - Score: 51.31
- [Grok 4.1 Fast](/models/grok-4-1-fast) - Score: 47.26
- [Grok 3 Mini](/models/grok-3-mini) - Score: 46.52
- [Grok Code Fast 1](/models/grok-code-fast-1) - Score: 34.16
- [Grok Build 0.1](/models/grok-build-0-1) - Score: not computed
- [Grok TTS](/models/grok-tts) - Score: not computed
- [Grok 4.20 Multi-agent](/models/grok-4-20-multi-agent-beta) - Score: not computed
- [Grok Realtime](/models/grok-realtime) - Score: not computed
- [Grok Voice Think Fast 1.0](/models/grok-voice-think-fast-1-0) - Score: not computed
- [Grok Voice Think Fast 2.0](/models/grok-voice-think-fast-2-0) - Score: not computed
