# DeepSeek V4 Flash 0731 Benchmark Scores & Performance

> DeepSeek V4 Flash 0731 by DeepSeek has 54 source-displayable benchmark rows but no public overall score. It remains unranked.

Canonical page: https://benchlm.ai/models/deepseek-v4-flash-0731

Last updated: September 10, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | DeepSeek |
| Source Type | Proprietary |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Overall Score | Coming soon |
| Overall Rank | Unranked |

## Family & Coverage

- Family: DeepSeek V4 Flash
- Variant: flash-reasoning (0731)
- Benchmarks covered: 54 of 434
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 56.9% |
| [BrowseComp](/benchmarks/browsecomp) | 73.2% |
| [HLE w/ tools](/benchmarks/hlewithtools) | 45.1% |
| [MCP Atlas](/benchmarks/mcpatlas) | 69% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 1189 |
| [Toolathlon](/benchmarks/toolathlon) | 47.8% |
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 82.7% |
| [CyberGym](/benchmarks/cybergym) | 76.7% |
| [Toolathlon-Verified](/benchmarks/toolathlonverified) | 70.3% |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 25.2% |
| [AutomationBench](/benchmarks/automationbench) | 25.1% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 48.4% |
| [Terminal-Bench 2.1 (Vals)](/benchmarks/valsterminalbench21) | 67.0% |
| [AA Agentic Index](/benchmarks/aaagenticindex) | 41.7% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [LiveCodeBench Pass@1-COT](/benchmarks/livecodebenchpass1cot) | 91.6% |
| [Codeforces](/benchmarks/codeforces) | 3052.0 |
| [SWE-bench Verified](/benchmarks/swe-bench-verified) | 79% |
| [SWE-bench Pro](/benchmarks/swe-bench-pro) | 52.6% |
| [SWE Multilingual](/benchmarks/swe-bench-multilingual) | 73.3% |
| [Terminal-Bench 2.0](/benchmarks/terminal-bench-2) | 56.9% |
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 82.7% |
| [NL2Repo](/benchmarks/nl2repo) | 54.2% |
| [DeepSWE](/benchmarks/deepswe) | 54.4% |
| [DSBench-FullStack](/benchmarks/dsbenchfullstack) | 68.7% |
| [DSBench-Hard](/benchmarks/dsbenchhard) | 59.6% |
| [AA-SciCode](/benchmarks/aascicode) | 50.3% |
| [VulcanBench v3](/benchmarks/vulcanbench) | 88.4% |
| [OpenHarmony Bench](/benchmarks/openharmonybench) | 53.8% |
| [LiveCodeBench (Vals)](/benchmarks/valslivecodebench) | 87.3% |
| [SWE-bench (Vals)](/benchmarks/valsswebench) | 88.8% |
| [AA Coding Index](/benchmarks/aacodingindex) | 69.1% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Design Arena Website](/benchmarks/designarenawebsite) | 1220 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MRCR 1M](/benchmarks/mrcr1m) | 78.7% |
| [CorpusQA 1M](/benchmarks/corpusqa1m) | 60.5% |
| [AA-LCR](/benchmarks/lcr) | 79.7% |
| [CritPt](/benchmarks/critpt) | 16.6% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [MMLU-Pro](/benchmarks/mmlu-pro) | 86.2% |
| [SimpleQA](/benchmarks/simpleqa) | 34.1% |
| [Chinese-SimpleQA](/benchmarks/chinesesimpleqa) | 78.9% |
| [GPQA](/benchmarks/gpqa) | 88.1% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 88.1% |
| [HLE](/benchmarks/hle) | 34.8% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 34.5% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 90.8% |
| [AA-HLE](/benchmarks/aahle) | 38.6% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | -14.3% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 40.4% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 91.7% |
| [GPQA Diamond (Vals)](/benchmarks/valsgpqadiamond) | 89.9% |
| [MMLU-Pro (Vals)](/benchmarks/valsmmlupro) | 86.2% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [HMMT Feb 2026](/benchmarks/hmmtfeb2026) | 94.8% |
| [IMOAnswerBench](/benchmarks/imoanswerbench) | 88.4% |
| [Apex](/benchmarks/apex) | 33.0% |
| [Apex Shortlist](/benchmarks/apexshortlist) | 85.7% |

## Other DeepSeek Models

- [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) - Score: 66.38
- [DeepSeek V3.2](/models/deepseek-v3-2) - Score: 56.88
- [DeepSeek V3.2 (Thinking)](/models/deepseek-v3-2-thinking) - Score: 55.8
- [DeepSeek LLM 2.0](/models/deepseek-llm-2-0) - Score: 52.26
- [DeepSeek V3.1](/models/deepseek-v3-1) - Score: 50.61
- [DeepSeek-R1](/models/deepseek-r1) - Score: 50.21
- [DeepSeek V3.1 (Reasoning)](/models/deepseek-v3-1-reasoning) - Score: 48.8
- [DeepSeek Coder 2.0](/models/deepseek-coder-2-0) - Score: 48.09
- [DeepSeekMath V2](/models/deepseekmath-v2) - Score: 47.71
- [DeepSeek V3](/models/deepseek-v3) - Score: 41.49
- [DeepSeek R1 Distill Qwen 32B](/models/deepseek-r1-distill-qwen-32b) - Score: 32.45
- [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) - Score: not computed
