# DeepSeek V3 Benchmark Scores & Performance

> DeepSeek V3 by DeepSeek scores 43.74/100 overall, ranking #168 out of 411 AI models.

Canonical page: https://benchlm.ai/models/deepseek-v3

Last updated: September 4, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | DeepSeek |
| Source Type | Open Weight |
| Reasoning Type | Non-Reasoning |
| Context Window | 128K |
| Overall Score | 43.74/100 |
| Overall Rank | #168 of 411 |

## Family & Coverage

- Family: DeepSeek
- Variant: snapshot (V3)
- Benchmarks covered: 22 of 422
- Coverage note: BenchLM currently has partial benchmark coverage for this model, so the overall score is conservative.

## Agentic Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA Agentic Index](/benchmarks/aaagenticindex) | 1.6% |
| [τ²-bench results](/benchmarks/tau2-bench) | 22.8% |
| [GDPval-AA](/benchmarks/gdpvalaanormalized) | 0.0% |
| [GDPval-AA](/benchmarks/gdpvalaa) | 235 |

## Coding Benchmarks

| Benchmark | Score |
|-----------|-------|
| [LiveCodeBench](/benchmarks/livecodebench) | 37.6% |
| [SWE-bench Verified](/benchmarks/swe-bench-verified) | 42% |
| [AA Coding Index](/benchmarks/aacodingindex) | 23.0% |
| [AA-SciCode](/benchmarks/aascicode) | 35.4% |

## Multimodal & Grounded Benchmarks

| Benchmark | Score |
|-----------|-------|
| [Design Arena Website](/benchmarks/designarenawebsite) | 1135 |

## Reasoning Benchmarks

| Benchmark | Score |
|-----------|-------|
| [AA-LCR](/benchmarks/lcr) | 31.7% |
| [CritPt](/benchmarks/critpt) | 0.0% |

## Knowledge Benchmarks

| Benchmark | Score |
|-----------|-------|
| [GPQA](/benchmarks/gpqa) | 59.1% |
| [MMLU-Pro](/benchmarks/mmlu-pro) | 75.9% |
| [Artificial Analysis Intelligence Index](/benchmarks/artificialanalysis) | 14.2% |
| [AA-GPQA Diamond](/benchmarks/aagpqadiamond) | 55.7% |
| [AA-HLE](/benchmarks/aahle) | 2.9% |
| [AA-Omniscience Index](/benchmarks/aaomniscienceindex) | -41.6% |
| [AA-Omniscience Accuracy](/benchmarks/omniscienceaccuracy) | 25.5% |
| [AA-Omniscience Hallucination Rate](/benchmarks/omnisciencehallucinationrate) | 90.0% |

## Instruction Following Benchmarks

| Benchmark | Score |
|-----------|-------|
| [IFEval](/benchmarks/ifeval) | 86.1% |
| [AA-IFBench](/benchmarks/aaifbench) | 34.8% |

## Mathematics Benchmarks

| Benchmark | Score |
|-----------|-------|
| [FrontierMath v2 (Tiers 1-3)](/benchmarks/frontiermathv2tiers13) | 1.724% |

## Other DeepSeek Models

- [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) - Score: 66.35
- [DeepSeek V3.2](/models/deepseek-v3-2) - Score: 57.62
- [DeepSeek V3.2 (Thinking)](/models/deepseek-v3-2-thinking) - Score: 57.54
- [DeepSeek LLM 2.0](/models/deepseek-llm-2-0) - Score: 54.01
- [DeepSeek V3.1](/models/deepseek-v3-1) - Score: 52.67
- [DeepSeek V3.1 (Reasoning)](/models/deepseek-v3-1-reasoning) - Score: 51.76
- [DeepSeek-R1](/models/deepseek-r1) - Score: 51.33
- [DeepSeek Coder 2.0](/models/deepseek-coder-2-0) - Score: 49.84
- [DeepSeekMath V2](/models/deepseekmath-v2) - Score: 49.45
- [DeepSeek R1 Distill Qwen 32B](/models/deepseek-r1-distill-qwen-32b) - Score: 34.18
- [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) - Score: not computed
