# DeepSeek V4.1 Flash Benchmark Scores & Performance

> DeepSeek V4.1 Flash by DeepSeek has 22 source-displayable benchmark rows but no public overall score. It remains unranked.

Canonical page: https://benchlm.ai/models/deepseek-v4-1-flash

Last updated: September 10, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | DeepSeek |
| Source Type | Open Weight |
| Reasoning Type | Reasoning |
| Context Window | 1M |
| Official model card | [DeepSeek-V4.1-Flash model card](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash) |
| Overall Score | Coming soon |
| Overall Rank | Unranked |

## Family & Coverage

- Family: DeepSeek V4.1 Flash
- Variant: flash-reasoning
- Benchmarks covered: 22 of 428
- Related earlier model: [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731)
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 90.6% |
| [terminalBench3](/benchmarks/terminalbench3) | 30% |
| [Terminal-Bench 4.0](/benchmarks/terminal-bench-4) | 31.20% |
| [CyberGym](/benchmarks/cybergym) | 88.1% |
| [ExploitGym](/benchmarks/exploitgym) | 15.3% |
| [HLE w/ tools](/benchmarks/hlewithtools) | 63.9% |
| [AutomationBench](/benchmarks/automationbench) | 54.8% |
| [Agents' Last Exam](/benchmarks/agentslastexam) | 31.8% |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Codeforces](/benchmarks/codeforces) | 3471.0 |
| [Terminal-Bench 2.1](/benchmarks/terminalbench21) | 90.6% |
| [terminalBench3](/benchmarks/terminalbench3) | 30% |
| [DeepSWE](/benchmarks/deepswe) | 74.2% |
| [ProgramBench](/benchmarks/programbench) | 20.3% |
| [NL2Repo](/benchmarks/nl2repo) | 65.4% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Chartography (tools)](/benchmarks/chartographywithtools) | 78.9% |
| [BabyVision w/ Python](/benchmarks/babyvisionpython) | 89.6% |
| [ZeroBench w/ Python](/benchmarks/zerobenchpython) | 49.0% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 90.9% |
| [GPQA-D](/benchmarks/gpqa-diamond) | 90.9% |
| [HLE](/benchmarks/hle) | 36.8% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Apex](/benchmarks/apex) | 65.6% |

## Other DeepSeek Models

- [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) - Score: 66.4
- [DeepSeek V3.2](/models/deepseek-v3-2) - Score: 56.91
- [DeepSeek V3.2 (Thinking)](/models/deepseek-v3-2-thinking) - Score: 56.34
- [DeepSeek LLM 2.0](/models/deepseek-llm-2-0) - Score: 52.81
- [DeepSeek V3.1](/models/deepseek-v3-1) - Score: 50.64
- [DeepSeek-R1](/models/deepseek-r1) - Score: 50.24
- [DeepSeek V3.1 (Reasoning)](/models/deepseek-v3-1-reasoning) - Score: 48.82
- [DeepSeek Coder 2.0](/models/deepseek-coder-2-0) - Score: 48.64
- [DeepSeekMath V2](/models/deepseekmath-v2) - Score: 48.25
- [DeepSeek V3](/models/deepseek-v3) - Score: 41.54
- [DeepSeek R1 Distill Qwen 32B](/models/deepseek-r1-distill-qwen-32b) - Score: 32.62
- [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) - Score: not computed
