# Claude Haiku 4.5 vs Grok 4.20

> Side-by-side benchmark comparison for Claude Haiku 4.5 and Grok 4.20 across agentic, coding, multimodal, knowledge, reasoning, multilingual, and math tasks.

- Canonical page: https://benchlm.ai/compare/claude-haiku-4-5-vs-grok-4-20-beta
- Last updated: September 10, 2026

- Shared sourced benchmarks: 6
- HTML indexing: indexable
- Ranking lane: BenchAlign v5

## Quick Verdict

Pick Grok 4.20 if you want the stronger benchmark profile. Claude Haiku 4.5 only makes more sense when its price, context window, or workload-specific category wins matter more than the overall score.

## Summary

- Grok 4.20 leads overall 67.13 to 52.79.
- The clearest category separation is in knowledge, where the averages are 44.2 for Claude Haiku 4.5 and 49.4 for Grok 4.20.
- The biggest single benchmark swing is LiveCodeBench (Vals) in Coding, with scores of 41.2% and 84.3%.
- Claude Haiku 4.5 is the cheaper option on output tokens, which matters if you expect large responses or heavy interactive use.
- Grok 4.20 also has the larger context window at 2M.

## Model Snapshot

| Property | Claude Haiku 4.5 | Grok 4.20 |
|----------|----------|----------|
| Creator | Anthropic | xAI |
| Type | Proprietary | Proprietary |
| Reasoning | Non-Reasoning | Reasoning |
| Context | 200K | 2M |
| Overall Score | 52.79 | 67.13 |
| Benchmarks Covered | 10 | 23 |
| Pricing (input/output) | $1.00 / $5.00 | $2.00 / $6.00 |

## Category Breakdown

### Agentic

- Winner: Tie
- Claude Haiku 4.5 public-lane score: 27 (Supported · #142/152)
- Grok 4.20 public-lane score: 26.7 (Supported · #145/152)

| Benchmark | Claude Haiku 4.5 | Grok 4.20 | Winner |
|-----------|-----------|-----------|--------|
| JobBench | 16.0% | Coming soon | Coming soon |
| Terminal-Bench 2.1 (Vals) | 43.8% | 44.2% | Grok 4.20 |
| Terminal-Bench 2.0 | Coming soon | 47.1% | Coming soon |
| DeepSearchQA | Coming soon | 62.8% | Coming soon |
| Gert Labs | Coming soon | 38.36% | Coming soon |

### Coding

- Winner: Grok 4.20
- Claude Haiku 4.5 public-lane score: 27.2 (Supported · #144/151)
- Grok 4.20 public-lane score: 28.2 (Supported · #141/151)

| Benchmark | Claude Haiku 4.5 | Grok 4.20 | Winner |
|-----------|-----------|-----------|--------|
| SWE-bench Verified | 73.3% | 76.7% | Grok 4.20 |
| VulcanBench v3 | 76.2% | Coming soon | Coming soon |
| LiveCodeBench (Vals) | 41.2% | 84.3% | Grok 4.20 |
| SWE-bench (Vals) | 66.6% | 72.2% | Grok 4.20 |
| LiveCodeBench Pro | Coming soon | 74.2% | Coming soon |
| SWE-bench Pro | Coming soon | 51.8% | Coming soon |
| Vibe Code Bench | Coming soon | 4.06% | Coming soon |

### Multimodal & Grounded

- Winner: Coming soon
- Claude Haiku 4.5 public-lane score: Coming soon
- Grok 4.20 public-lane score: 34.6 (#43/48)

| Benchmark | Claude Haiku 4.5 | Grok 4.20 | Winner |
|-----------|-----------|-----------|--------|
| MMMU-Pro | Coming soon | 75.2% | Coming soon |
| CharXiv | Coming soon | 60.9% | Coming soon |
| ERQA | Coming soon | 54.1% | Coming soon |
| SimpleVQA | Coming soon | 57.4% | Coming soon |
| MedXpertQA (MM) | Coming soon | 65.8% | Coming soon |

### Reasoning

- Winner: Coming soon
- Claude Haiku 4.5 public-lane score: Coming soon
- Grok 4.20 public-lane score: 34.2 (Unranked · 2 rankable rows)

| Benchmark | Claude Haiku 4.5 | Grok 4.20 | Winner |
|-----------|-----------|-----------|--------|
| ARC-AGI-2 | Coming soon | 53.3% | Coming soon |
| ARC-AGI-3 | Coming soon | 0.1% | Coming soon |

### Knowledge

- Winner: Coming soon
- Claude Haiku 4.5 public-lane score: 44.2 (Estimated · #115/183)
- Grok 4.20 public-lane score: 49.4 (Supported · #85/183)

| Benchmark | Claude Haiku 4.5 | Grok 4.20 | Winner |
|-----------|-----------|-----------|--------|
| GPQA Diamond (Vals) | 72.2% | 88.6% | Grok 4.20 |
| MMLU-Pro (Vals) | 78.7% | 86.3% | Grok 4.20 |
| GPQA-D | Coming soon | 88.5% | Coming soon |
| HLE w/o tools | Coming soon | 31.6% | Coming soon |
| HealthBench Hard | Coming soon | 20.3% | Coming soon |
| MedXpertQA (Text) | Coming soon | 50.2% | Coming soon |

### Mathematics

- Winner: Coming soon
- Claude Haiku 4.5 public-lane score: 28.9 (Unranked · 2 rankable rows)
- Grok 4.20 public-lane score: Coming soon

| Benchmark | Claude Haiku 4.5 | Grok 4.20 | Winner |
|-----------|-----------|-----------|--------|
| FrontierMath v2 (Tiers 1-3) | 5.903% | Coming soon | Coming soon |
| FrontierMath v2 (Tier 4) | 2.083% | Coming soon | Coming soon |

## FAQ

### Which is better overall, Claude Haiku 4.5 or Grok 4.20?

Grok 4.20 is ahead overall on BenchLM right now.

### Where is the biggest gap between Claude Haiku 4.5 and Grok 4.20?

The widest category gap is in knowledge, where the averages are 44.2 for Claude Haiku 4.5 and 49.4 for Grok 4.20.

### How many benchmarks does Claude Haiku 4.5 cover on BenchLM?

Claude Haiku 4.5 currently has 10 sourced benchmark scores on BenchLM.

### How many benchmarks does Grok 4.20 cover on BenchLM?

Grok 4.20 currently has 23 sourced benchmark scores on BenchLM.

## Related Comparisons

- [Claude Haiku 4.5 vs Claude Haiku 4.5 Thinking](/compare/claude-haiku-4-5-vs-claude-haiku-4-5-thinking)
- [Grok 4.20 vs Claude Haiku 4.5 Thinking](/compare/claude-haiku-4-5-thinking-vs-grok-4-20-beta)
- [Claude Haiku 4.5 vs Grok 4.20 Multi-agent](/compare/claude-haiku-4-5-vs-grok-4-20-multi-agent-beta)
- [Grok 4.20 vs Grok 4.20 Multi-agent](/compare/grok-4-20-beta-vs-grok-4-20-multi-agent-beta)

## Explore More

- [Claude Haiku 4.5 profile](/models/claude-haiku-4-5)
- [Grok 4.20 profile](/models/grok-4-20-beta)
- [Compare Pricing](/llm-pricing)
- [Alternative Finder](/tools/alternative-finder)
- [LLM Selector](/tools/llm-selector)
- [Overall Rankings](/best/overall)
