# Grok 4.20 vs Hy3 Preview

> Side-by-side benchmark comparison for Grok 4.20 and Hy3 Preview across agentic, coding, multimodal, knowledge, reasoning, multilingual, and math tasks.

- Canonical page: https://benchlm.ai/compare/grok-4-20-beta-vs-hy3-preview
- Last updated: September 23, 2026

- Shared sourced benchmarks: 4
- HTML indexing: indexable
- Ranking lane: BenchAlign v5.6

## Quick Verdict

Pick Grok 4.20 if you want the stronger benchmark profile. Hy3 Preview only makes more sense when its price, context window, or workload-specific category wins matter more than the overall score.

## Summary

- Grok 4.20 leads overall 59.56 to 45.32.
- The clearest category separation is in coding, where the averages are 22.4 for Grok 4.20 and 33.3 for Hy3 Preview.
- The biggest single benchmark swing is Terminal-Bench 2.0 in Agentic, with scores of 47.1% and 54.4%.
- Hy3 Preview is the cheaper option on output tokens, which matters if you expect large responses or heavy interactive use.
- Grok 4.20 also has the larger context window at 2M.

## Model Snapshot

| Property | Grok 4.20 | Hy3 Preview |
|----------|----------|----------|
| Creator | xAI | Tencent |
| Type | Proprietary | Open Weight |
| Reasoning | Reasoning | Reasoning |
| Context | 2M | 256K |
| Overall Score | 59.56 | 45.32 |
| Benchmarks Covered | 23 | 6 |
| Pricing (input/output) | $2.00 / $6.00 | $0.00 / $0.00 |

## Category Breakdown

### Agentic

- Winner: Not comparable
- Grok 4.20 public-lane score: 22.3 (Supported · #83/105)
- Hy3 Preview public-lane score: Coming soon

| Benchmark | Grok 4.20 | Hy3 Preview | Winner |
|-----------|-----------|-----------|--------|
| Terminal-Bench 2.0 | 47.1% | 54.4% | Hy3 Preview |
| DeepSearchQA | 62.8% | Coming soon | Coming soon |
| Gert Labs | 38.36% | 36.91% | Grok 4.20 |
| Terminal-Bench 2.1 (Vals) | 44.2% | Coming soon | Coming soon |

### Coding

- Winner: Directional only
- Grok 4.20 public-lane score: 22.4 (Supported · #121/135)
- Hy3 Preview public-lane score: 33.3 (Estimated · #85/135)

| Benchmark | Grok 4.20 | Hy3 Preview | Winner |
|-----------|-----------|-----------|--------|
| LiveCodeBench Pro | 74.2% | Coming soon | Coming soon |
| SWE-bench Verified | 76.7% | 74.4% | Grok 4.20 |
| SWE-bench Pro | 51.8% | Coming soon | Coming soon |
| Vibe Code Bench | 4.06% | Coming soon | Coming soon |
| LiveCodeBench (Vals) | 84.3% | Coming soon | Coming soon |
| SWE-bench (Vals) | 72.2% | Coming soon | Coming soon |
| Terminal-Bench 2.0 | Coming soon | 54.4% | Coming soon |

### Reasoning

- Winner: Not comparable
- Grok 4.20 public-lane score: 34.3 (Unranked · 2 rankable rows)
- Hy3 Preview public-lane score: Coming soon

| Benchmark | Grok 4.20 | Hy3 Preview | Winner |
|-----------|-----------|-----------|--------|
| ARC-AGI-2 | 53.3% | Coming soon | Coming soon |
| ARC-AGI-3 | 0.1% | Coming soon | Coming soon |

### Multimodal & Grounded

- Winner: Not comparable
- Grok 4.20 public-lane score: 34.6 (#45/50)
- Hy3 Preview public-lane score: Coming soon

| Benchmark | Grok 4.20 | Hy3 Preview | Winner |
|-----------|-----------|-----------|--------|
| MMMU-Pro | 75.2% | Coming soon | Coming soon |
| CharXiv | 60.9% | Coming soon | Coming soon |
| ERQA | 54.1% | Coming soon | Coming soon |
| SimpleVQA | 57.4% | Coming soon | Coming soon |
| MedXpertQA (MM) | 65.8% | Coming soon | Coming soon |

### Knowledge

- Winner: Directional only
- Grok 4.20 public-lane score: 43.3 (Supported · #76/160)
- Hy3 Preview public-lane score: 39.1 (Estimated · #93/160)

| Benchmark | Grok 4.20 | Hy3 Preview | Winner |
|-----------|-----------|-----------|--------|
| GPQA-D | 88.5% | 87.2% | Grok 4.20 |
| HLE w/o tools | 31.6% | Coming soon | Coming soon |
| HealthBench Hard | 20.3% | Coming soon | Coming soon |
| MedXpertQA (Text) | 50.2% | Coming soon | Coming soon |
| GPQA Diamond (Vals) | 88.6% | Coming soon | Coming soon |
| MMLU-Pro (Vals) | 86.3% | Coming soon | Coming soon |
| GPQA | Coming soon | 87.2% | Coming soon |

### Instruction Following

- Winner: Not comparable
- Grok 4.20 public-lane score: Coming soon
- Hy3 Preview public-lane score: 46.9 (#86/124)

## FAQ

### Which is better overall, Grok 4.20 or Hy3 Preview?

Grok 4.20 is ahead overall on BenchLM right now.

### Where is the biggest gap between Grok 4.20 and Hy3 Preview?

The widest category gap is in coding, where the averages are 22.4 for Grok 4.20 and 33.3 for Hy3 Preview.

### How many benchmarks does Grok 4.20 cover on BenchLM?

Grok 4.20 currently has 23 sourced benchmark scores on BenchLM.

### How many benchmarks does Hy3 Preview cover on BenchLM?

Hy3 Preview currently has 6 sourced benchmark scores on BenchLM.

## Related Comparisons

- [Grok 4.20 vs Grok 4.20 Multi-agent](/compare/grok-4-20-beta-vs-grok-4-20-multi-agent-beta)
- [Hy3 Preview vs Grok 4.20 Multi-agent](/compare/grok-4-20-multi-agent-beta-vs-hy3-preview)
- [Grok 4.20 vs Hy3](/compare/grok-4-20-beta-vs-hy3)
- [Hy3 Preview vs Hy3](/compare/hy3-vs-hy3-preview)

## Explore More

- [Grok 4.20 profile](/models/grok-4-20-beta)
- [Hy3 Preview profile](/models/hy3-preview)
- [Compare Pricing](/llm-pricing)
- [Alternative Finder](/tools/alternative-finder)
- [LLM Selector](/tools/llm-selector)
- [Overall Rankings](/best/overall)
