# Step 5 Preview Benchmark Scores & Performance

> Step 5 Preview by StepFun has 34 source-displayable benchmark rows but no public overall score. It remains unranked.

Canonical page: https://benchlm.ai/models/step-5-preview

Last updated: September 21, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | StepFun |
| Source Type | Pending |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Official model card | [StepFun model documentation](https://platform.stepfun.ai/docs/en/guides/models/step-5-preview) |
| Overall Score | Coming soon |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Step 5
- Variant: preview (Preview)
- Benchmarks covered: 34 of 447
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 85.0% |
| [Terminal-Bench 4.0](/benchmarks/terminal-bench-4) | 33.30% |
| [CyberGym](/benchmarks/cybergym) | 84.7% |
| [AutomationBench](/benchmarks/automationbench) | 44.0% |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 74.1% |
| [MCP Atlas](/benchmarks/mcpatlas) | 85.6% |
| [JobBench](/benchmarks/jobbench) | 59.0% |
| [APEX-Agents](/benchmarks/apexagents) | 37.8% |
| [DRACO](/benchmarks/draco) | 83.3% |
| [BrowseComp](/benchmarks/browsecomp) | 88.7% |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 29.5% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1566 |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 53.3% |
| [AA Briefcase](/benchmarks/aabriefcaseelo) | 1432 |
| [GDP.pdf](/benchmarks/aagdppdf) | 14.8% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [DeepSWE](/benchmarks/deepswe) | 67.7% |
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 85.0% |
| [SciCode](/benchmarks/scicode) | 58.9% |
| [ProgramBench](/benchmarks/programbench) | 80.5% |
| [sweMarathon](/benchmarks/swemarathon) | 72.7% |
| [MLS-Bench Lite](/benchmarks/mlsbenchlite) | 40.5% |
| [AA-SciCode](/benchmarks/aascicode) | 58.9% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [OfficeQA Pro](/benchmarks/officeqapro) | 60.3% |
| [MMMU-Pro](/benchmarks/mmmu-pro) | 76% |
| [AA-MMMU-Pro](/benchmarks/aammmupro) | 76.4% |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 88.3% |
| [CritPt](/benchmarks/critpt) | 20.9% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA-D](/benchmarks/gpqa-diamond) | 93.5% |
| [HLE](/benchmarks/hle) | 46.5% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 43.7% |
| [AA-HLE](/benchmarks/aahle) | 46.5% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 16.4% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 41.5% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 43.0% |

## Other StepFun Models

- [Step 3.7 Flash](/models/step-3-7-flash) - Score: 49.93
- [Step 3.5 Flash](/models/step-3-5-flash) - Score: 49.83
- [Step-Audio](/models/step-audio) - Score: not computed
- [Step-Audio-Chat 130B](/models/step-audio-chat-130b) - Score: not computed
