# Hy4 preview Benchmark Scores & Performance

> Hy4 preview by Tencent scores 62.1/100 overall, ranking #51 out of 499 AI models.

Canonical page: https://benchlm.ai/models/hy4-preview

Last updated: September 21, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | Tencent |
| Source Type | Open Weight |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Official model card | [Tencent Hy4 preview model card](https://huggingface.co/tencent/Hy4-preview) |
| Overall Score | 62.1/100 |
| Overall Rank | #51 of 499 |

## Family & Coverage

- Family: Hy4
- Variant: preview (Preview)
- Benchmarks covered: 30 of 447
- Related earlier model: [Hy3](/models/hy3)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 85.4% |
| [CyberGym](/benchmarks/cybergym) | 78.4% |
| [WideResearch](/benchmarks/wideresearch) | 83.9% |
| [DRACO](/benchmarks/draco) | 77.2% |
| [MCP Atlas](/benchmarks/mcpatlas) | 83.7% |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 74.1% |
| [APEX-Agents](/benchmarks/apexagents) | 37.1% |
| [skillsBench](/benchmarks/skillsbench) | 62.9% |
| [JobBench](/benchmarks/jobbench) | 61.7% |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 22.8% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1678 |
| [AutomationBench](/benchmarks/automationbench) | 32.1% |
| [BankerToolBench](/benchmarks/bankertoolbench) | 78.6% |
| [HLE w/ tools](/benchmarks/hlewithtools) | 55.4% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 55.1% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 85.4% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 65.7% |
| [SWE Multilingual](/benchmarks/swe-bench-multilingual) | 82.9% |
| [DeepSWE](/benchmarks/deepswe) | 64.3% |
| [NL2Repo](/benchmarks/nl2repo) | 58.9% |
| [ProgramBench](/benchmarks/programbench) | 17.5% |
| [PostTrain Bench](/benchmarks/posttrainbench) | 35.6% |
| [sweMarathon](/benchmarks/swemarathon) | 31.9% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [OfficeQA Pro](/benchmarks/officeqapro) | 66.2% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [CritPt](/benchmarks/critpt) | 16.9% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 92.3% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 92.3% |
| [HLE](/benchmarks/hle) | 55.4% |
| [HLE w/o tools](/benchmarks/hlenotools) | 43.4% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Apex](/benchmarks/apex) | 74.2% |

## Other Tencent Models

- [Hy3](/models/hy3) - Score: 60.59
- [Hy3 Preview](/models/hy3-preview) - Score: 56.02
- [UI-Mate-27B](/models/ui-mate-27b) - Score: not computed
- [UI-Mate-9B](/models/ui-mate-9b) - Score: not computed
- [AuK](/models/tencent-auk) - Score: not computed
- [Hy ASR 3.0 preview](/models/hy-asr-3-0-preview) - Score: not computed
- [Hy-MT2-1.8B](/models/hy-mt2-1-8b) - Score: not computed
- [Hy-MT2-30B-A3B](/models/hy-mt2-30b-a3b) - Score: not computed
