# DeepSeek V4 Pro 0813 Benchmark Scores & Performance

> DeepSeek V4 Pro 0813 by DeepSeek scores 66.4/100 overall, ranking #32 out of 412 AI models.

Canonical page: https://benchlm.ai/models/deepseek-v4-pro-0813

Last updated: September 8, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | DeepSeek |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Overall Score | 66.4/100 |
| Overall Rank | #32 of 412 |

## Family & Coverage

- Family: DeepSeek V4 Pro
- Variant: pro-reasoning (0813)
- Benchmarks covered: 61 of 427
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 67.9% |
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 87.9% |
| [BrowseComp](/benchmarks/browsecomp) | 83.4% |
| [HLE w/ tools](/benchmarks/hlewithtools) | 60.0% |
| [MCP Atlas](/benchmarks/mcpatlas) | 73.6% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1306 |
| [Toolathlon](/benchmarks/toolathlon) | 51.8% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 49.6% |
| [APEX-Agents-AA](/benchmarks/apexagentsaa) | 24.3% |
| [τ²-bench results](/benchmarks/tau2-bench) | 96.2% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 54.5% |
| [CyberGym](/benchmarks/cybergym) | 83.3% |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 74.1% |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 25.7% |
| [AutomationBench](/benchmarks/automationbench) | 31.8% |
| [AA Tau3 Banking](/benchmarks/aatau3banking) | 39.6% |
| [AA EnterpriseOps-Gym](/benchmarks/aaenterpriseopsgym) | 49.6% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 54.7% |
| [AA Briefcase](/benchmarks/aabriefcaseelo) | 1265 |
| [AA AutomationBench](/benchmarks/aaautomationbench) | 56.7% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [LiveCodeBench Pass@1-COT](/benchmarks/livecodebenchpass1cot) | 93.5% |
| [Codeforces](/benchmarks/codeforces) | 3206.0 |
| [SWE-bench Verified](/benchmarks/swe-bench-verified) | 80.6% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 55.4% |
| [SWE Multilingual](/benchmarks/swe-bench-multilingual) | 76.2% |
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 67.9% |
| [Vibe Code Bench](/benchmarks/vibecodebench) | 49.93% |
| [AA Coding Index](/benchmarks/aacodingindex) | 68.8% |
| [AA-SciCode](/benchmarks/aascicode) | 51.0% |
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 87.9% |
| [NL2Repo](/benchmarks/nl2repo) | 61.5% |
| [deepSwe](/benchmarks/deepswe) | 62.7% |
| [DSBench-FullStack](/benchmarks/dsbenchfullstack) | 71.1% |
| [DSBench-Hard](/benchmarks/dsbenchhard) | 67.2% |
| [OpenHarmony Bench](/benchmarks/openharmonybench) | 59.0% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 87.5% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 96.4% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Design Arena Website](/benchmarks/designarenawebsite) | 1258 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MRCR 1M](/benchmarks/mrcr1m) | 83.5% |
| [CorpusQA 1M](/benchmarks/corpusqa1m) | 62.0% |
| [AA-LCR](/benchmarks/lcr) | 80.3% |
| [CritPt](/benchmarks/critpt) | 18.0% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMLU-Pro](/benchmarks/mmlu-pro) | 87.5% |
| [SimpleQA](/benchmarks/simpleqa) | 57.9% |
| [Chinese-SimpleQA](/benchmarks/chinesesimpleqa) | 84.4% |
| [GPQA](/benchmarks/gpqa) | 90.1% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 90.1% |
| [HLE](/benchmarks/hle) | 42.7% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 36.3% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 92.8% |
| [AA-HLE](/benchmarks/aahle) | 41.0% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | 0.8% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 49.1% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 94.1% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 92.4% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 87.0% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-IFBench](/benchmarks/aaifbench) | 76.5% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [HMMT Feb 2026](/benchmarks/hmmtfeb2026) | 95.2% |
| [IMOAnswerBench](/benchmarks/imoanswerbench) | 89.8% |
| [Apex](/benchmarks/apex) | 38.3% |
| [Apex Shortlist](/benchmarks/apexshortlist) | 90.2% |

## Other DeepSeek Models

- [DeepSeek V3.2](/models/deepseek-v3-2) - Score: 56.91
- [DeepSeek V3.2 (Thinking)](/models/deepseek-v3-2-thinking) - Score: 56.32
- [DeepSeek LLM 2.0](/models/deepseek-llm-2-0) - Score: 52.78
- [DeepSeek V3.1](/models/deepseek-v3-1) - Score: 50.64
- [DeepSeek-R1](/models/deepseek-r1) - Score: 50.24
- [DeepSeek V3.1 (Reasoning)](/models/deepseek-v3-1-reasoning) - Score: 48.82
- [DeepSeek Coder 2.0](/models/deepseek-coder-2-0) - Score: 48.62
- [DeepSeekMath V2](/models/deepseekmath-v2) - Score: 48.23
- [DeepSeek V3](/models/deepseek-v3) - Score: 41.54
- [DeepSeek R1 Distill Qwen 32B](/models/deepseek-r1-distill-qwen-32b) - Score: 32.62
- [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) - Score: not computed
