# Grok 4.6 Benchmark Scores & Performance

> Grok 4.6 by xAI scores 69.19/100 overall, ranking #15 out of 507 AI models.

Canonical page: https://benchlm.ai/models/grok-4-6

Last updated: September 25, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | xAI |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 500K |
| Official model card | [xAI Grok 4.6 model documentation](https://docs.x.ai/developers/grok-4-6) |
| Overall Score | 69.19/100 |
| Overall Rank | #15 of 507 |

## Family & Coverage

- Family: Grok 4.6
- Variant: base
- Benchmarks covered: 37 of 486
- Related earlier model: [Grok 4.5](/models/grok-4-5)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 3.0](/benchmarks/terminal-bench-3) | 26.5% |
| [APEX-Agents](/benchmarks/apexagents) | 57.5% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 55.3% |
| [AA Briefcase](/benchmarks/aabriefcaseelo) | 1546 |
| [AA Tau3 Banking](/benchmarks/aatau3banking) | 50.7% |
| [AA EnterpriseOps-Gym](/benchmarks/aaenterpriseopsgym) | 48.3% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 78.3% |
| [AA AutomationBench](/benchmarks/aaautomationbench) | 66.7% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1605 |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 53.4% |
| [GDP.pdf](/benchmarks/aagdppdf) | 17.0% |
| [AA-AnalystAgent](/benchmarks/aaanalystagent) | 41.3% |
| [ApprenticeBench](/benchmarks/apprenticebench) | 13% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Bug Hunt Bench](/benchmarks/bug-hunt-bench) | 27 fixes |
| [DeepSWE](/benchmarks/deepswe) | 65.9% |
| [CursorBench 3.2](/benchmarks/cursorbench32) | 70.8% |
| [FrontierCode 1.1 Extended](/benchmarks/frontiercode11extended) | 61.3% |
| [AA-SciCode](/benchmarks/aascicode) | 56.5% |
| [VulcanBench v3](/benchmarks/vulcanbench) | 87.0% |
| [FrontierSWE v2](/benchmarks/frontierswev2) | 25.3% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 88.2% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 95.6% |
| [AA Coding Index](/benchmarks/aacodingindex) | 76.8% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Design Arena Website](/benchmarks/designarenawebsite) | 1302 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 80.3% |
| [CritPt](/benchmarks/critpt) | 17.1% |
| [ARC-AGI-1](/benchmarks/arcagi1) | 87.00% |
| [ARC-AGI-2](/benchmarks/arc-agi-2) | 67.1% |
| [ARC-AGI-3](/benchmarks/arcagi3) | 2.1% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 44.3% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 94.9% |
| [AA-HLE](/benchmarks/aahle) | 42.9% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 30.5% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 48.2% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 34.3% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 94.7% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 89.4% |

## Other xAI Models

- [Grok 4.5](/models/grok-4-5) - Score: 65.27
- [Grok 4.20](/models/grok-4-20-beta) - Score: 59.66
- [Grok 4.3](/models/grok-4-3) - Score: 54.73
- [Grok 4](/models/grok-4) - Score: 52.72
- [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) - Score: 45.68
- [Grok 3 [Beta]](/models/grok-3-beta) - Score: 45.04
- [Grok 3 Mini](/models/grok-3-mini) - Score: 43.64
- [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) - Score: 42.43
- [Grok 4.1 Fast](/models/grok-4-1-fast) - Score: 36.33
- [Grok Code Fast 1](/models/grok-code-fast-1) - Score: 32.05
- [Grok 4.7](/models/grok-4-7) - Score: not computed
- [Grok 4.1](/models/grok-4-1) - Score: not computed
- [Grok Voice Transcribe 2.0](/models/grok-voice-transcribe-2) - Score: not computed
- [Grok Build 0.1](/models/grok-build-0-1) - Score: not computed
- [Grok TTS](/models/grok-tts) - Score: not computed
- [Grok 4.20 Multi-agent](/models/grok-4-20-multi-agent-beta) - Score: not computed
- [Grok Realtime](/models/grok-realtime) - Score: not computed
- [Grok Voice Think Fast 1.0](/models/grok-voice-think-fast-1-0) - Score: not computed
- [Grok Voice Think Fast 2.0](/models/grok-voice-think-fast-2-0) - Score: not computed
