# DeepSeek V4 Flash 0731 vs Raw Qwen3 4B Instruct 2507 direct logits

> Side-by-side benchmark comparison for DeepSeek V4 Flash 0731 and Raw Qwen3 4B Instruct 2507 direct logits across agentic, coding, multimodal, knowledge, reasoning, multilingual, and math tasks.

- Canonical page: https://benchlm.ai/compare/deepseek-v4-flash-0731-vs-raw-qwen3-4b-instruct-2507
- Last updated: October 1, 2026

- Shared sourced benchmarks: 0
- HTML indexing: indexable
- Ranking lane: BenchAlign v5.8

## Quick Verdict

No benchmark verdict yet: there is no shared sourced benchmark row for DeepSeek V4 Flash 0731 and Raw Qwen3 4B Instruct 2507 direct logits.

## Summary

- DeepSeek V4 Flash 0731 and Raw Qwen3 4B Instruct 2507 direct logits do not currently share a sourced benchmark row. The mirror therefore compares metadata, pricing, context windows, and each model's separately reported coverage without naming a benchmark winner.

## Model Snapshot

| Property | DeepSeek V4 Flash 0731 | Raw Qwen3 4B Instruct 2507 direct logits |
|----------|----------|----------|
| Creator | DeepSeek | Alibaba |
| Type | Open Weight | Open Weight |
| Reasoning | Reasoning | Non-Reasoning |
| Context | 1M | null |
| Independent public score (not head-to-head) | null | null |
| Benchmarks Covered | 42 | 2 |
| Pricing (input/output) | $0.14 / $0.28 | Pricing unavailable / Pricing unavailable |

## Category Breakdown

### Agentic

- Winner: Insufficient shared evidence
- DeepSeek V4 Flash 0731 public-lane score: Coming soon
- Raw Qwen3 4B Instruct 2507 direct logits public-lane score: Coming soon

| Benchmark | DeepSeek V4 Flash 0731 | Raw Qwen3 4B Instruct 2507 direct logits | Winner |
|-----------|-----------|-----------|--------|
| Terminal-Bench 2.0 | 56.9% | Coming soon | Coming soon |
| BrowseComp | 73.2% | Coming soon | Coming soon |
| HLE w/ tools | 45.1% | Coming soon | Coming soon |
| MCP Atlas | 69% | Coming soon | Coming soon |
| Toolathlon | 47.8% | Coming soon | Coming soon |
| Terminal-Bench 2.1 | 82.7% | Coming soon | Coming soon |
| CyberGym | 76.7% | Coming soon | Coming soon |
| Toolathlon-Verified | 70.3% | Coming soon | Coming soon |
| Agents' Last Exam | 25.2% | Coming soon | Coming soon |
| AutomationBench | 25.1% | Coming soon | Coming soon |
| Terminal-Bench 2.1 (Vals) | 67.0% | Coming soon | Coming soon |

### Coding

- Winner: Insufficient shared evidence
- DeepSeek V4 Flash 0731 public-lane score: Coming soon
- Raw Qwen3 4B Instruct 2507 direct logits public-lane score: Coming soon

| Benchmark | DeepSeek V4 Flash 0731 | Raw Qwen3 4B Instruct 2507 direct logits | Winner |
|-----------|-----------|-----------|--------|
| LiveCodeBench Pass@1-COT | 91.6% | Coming soon | Coming soon |
| Codeforces | 3052.0 | Coming soon | Coming soon |
| SWE-bench Verified | 79% | Coming soon | Coming soon |
| SWE-bench Pro | 52.6% | Coming soon | Coming soon |
| SWE Multilingual | 73.3% | Coming soon | Coming soon |
| Terminal-Bench 2.0 | 56.9% | Coming soon | Coming soon |
| Terminal-Bench 2.1 | 82.7% | Coming soon | Coming soon |
| NL2Repo | 54.2% | Coming soon | Coming soon |
| DeepSWE | 54.4% | Coming soon | Coming soon |
| DSBench-FullStack | 68.7% | Coming soon | Coming soon |
| DSBench-Hard | 59.6% | Coming soon | Coming soon |
| VulcanBench v3 | 88.4% | Coming soon | Coming soon |
| OpenHarmony Bench | 53.8% | Coming soon | Coming soon |
| LiveCodeBench (Vals) | 87.3% | Coming soon | Coming soon |
| SWE-bench (Vals) | 88.8% | Coming soon | Coming soon |

### Reasoning

- Winner: Insufficient shared evidence
- DeepSeek V4 Flash 0731 public-lane score: 57 (Unranked · 4 rankable rows)
- Raw Qwen3 4B Instruct 2507 direct logits public-lane score: Coming soon

| Benchmark | DeepSeek V4 Flash 0731 | Raw Qwen3 4B Instruct 2507 direct logits | Winner |
|-----------|-----------|-----------|--------|
| MRCR 1M | 78.7% | Coming soon | Coming soon |
| CorpusQA 1M | 60.5% | Coming soon | Coming soon |
| ARC-AGI-1 | 89.00% | Coming soon | Coming soon |
| ARC-AGI-2 | 61.4% | Coming soon | Coming soon |
| JevBench 1.4 | Coming soon | 40.95 | Coming soon |
| JevBench 1.5 | Coming soon | 62.14 | Coming soon |

### Knowledge

- Winner: Insufficient shared evidence
- DeepSeek V4 Flash 0731 public-lane score: Coming soon
- Raw Qwen3 4B Instruct 2507 direct logits public-lane score: Coming soon

| Benchmark | DeepSeek V4 Flash 0731 | Raw Qwen3 4B Instruct 2507 direct logits | Winner |
|-----------|-----------|-----------|--------|
| MMLU-Pro | 86.2% | Coming soon | Coming soon |
| SimpleQA | 34.1% | Coming soon | Coming soon |
| Chinese-SimpleQA | 78.9% | Coming soon | Coming soon |
| GPQA | 88.1% | Coming soon | Coming soon |
| GPQA-D | 88.1% | Coming soon | Coming soon |
| HLE | 34.8% | Coming soon | Coming soon |
| GPQA Diamond (Vals) | 89.9% | Coming soon | Coming soon |
| MMLU-Pro (Vals) | 86.2% | Coming soon | Coming soon |

### Mathematics

- Winner: Insufficient shared evidence
- DeepSeek V4 Flash 0731 public-lane score: 79.9 (Unranked · 4 rankable rows)
- Raw Qwen3 4B Instruct 2507 direct logits public-lane score: Coming soon

| Benchmark | DeepSeek V4 Flash 0731 | Raw Qwen3 4B Instruct 2507 direct logits | Winner |
|-----------|-----------|-----------|--------|
| HMMT Feb 2026 | 94.8% | Coming soon | Coming soon |
| IMOAnswerBench | 88.4% | Coming soon | Coming soon |
| Apex | 33.0% | Coming soon | Coming soon |
| Apex Shortlist | 85.7% | Coming soon | Coming soon |

## FAQ

### Can I compare DeepSeek V4 Flash 0731 and Raw Qwen3 4B Instruct 2507 direct logits on BenchLM yet?

Not fully yet. BenchLM is tracking both models, but sourced benchmark coverage is still incomplete for a fair score-level comparison.

### Why does this page show "coming soon" values?

BenchLM only calls winners when public benchmark coverage is available on both sides of the comparison.

## Related Comparisons

- [DeepSeek V4 Flash 0731 vs GPT-6 Astra](/compare/deepseek-v4-flash-0731-vs-gpt-6-astra)
- [Raw Qwen3 4B Instruct 2507 direct logits vs GPT-6 Astra](/compare/gpt-6-astra-vs-raw-qwen3-4b-instruct-2507)
- [DeepSeek V4 Flash 0731 vs Claude Opus 5.5](/compare/claude-opus-5-5-vs-deepseek-v4-flash-0731)
- [Raw Qwen3 4B Instruct 2507 direct logits vs Claude Opus 5.5](/compare/claude-opus-5-5-vs-raw-qwen3-4b-instruct-2507)

## Explore More

- [DeepSeek V4 Flash 0731 profile](/models/deepseek-v4-flash-0731)
- [Raw Qwen3 4B Instruct 2507 direct logits profile](/models/raw-qwen3-4b-instruct-2507)
- [Compare Pricing](/llm-pricing)
- [Alternative Finder](/tools/alternative-finder)
- [LLM Selector](/tools/llm-selector)
- [Overall Rankings](/best/overall)
