Skip to main content
BenchLM ResearchDeepSeek V4 Flash

DeepSeek V4 Flash Gained 20.9 Points, Not 25.8

DeepSeek-V4-Flash-0731 shipped July 31 with MIT weights and $0.28 output pricing. The Terminal-Bench jump being quoted compares two benchmark versions, and DeepSeek's own table puts it behind Opus 4.8 on all nine rows.

Published
Last updated
Reading time
7 min
External sources
4
Tags: DeepSeek V4 Flash, DeepSeek, Terminal-Bench, open weightsData and scoring methodology
In this article6 sections

DeepSeek published DeepSeek-V4-Flash-0731 on July 31, 2026. Weights went to Hugging Face under the MIT License, and API rates held at $0.14 input and $0.28 output per million tokens.

One number has travelled faster than any of that: a 25.8-point jump on Terminal-Bench. DeepSeek's own model card does not report it.

That card puts 0731 at 82.7 on Terminal-Bench 2.1, and it puts the April Preview at 61.8 on the same test. Call it 20.9 points.

Where does 25.8 come from? Subtract 56.9, which is what Preview scored on Terminal-Bench 2.0, from 82.7, which is what 0731 scored on 2.1.

Twenty-point-nine is still a large single-checkpoint agentic gain, so nothing here is deflation. Version hygiene matters anyway. It separates a measured result from a number nobody ran.

DeepSeek shipped the weights, not just the endpoint

Anyone already pointed at the deepseek-v4-flash API ID is now served 0731 in public beta, with no migration and no new model string. DeepSeek calls it the official release superseding the preview, sharing its structure with DeepSeek-V4-Flash-DSpark, which means a speculative-decoding module ships attached to the checkpoint.

That module explains a discrepancy worth knowing before you quote a size. Hugging Face lists 304B parameters. Most coverage says 284B total and 13B activated, which described Preview.

Context length holds at 1M. reasoning_effort now accepts three levels — low, high, and max — and DeepSeek recommends temperature 1.0 with top_p 0.95 for agentic work, plus a 384K output ceiling at the two upper levels.

Pricing did not move: $0.14 per million cache-miss input tokens, $0.0028 on cache hits, $0.28 output. DeepSeek has said every billing item will cost double inside two Beijing-time windows each day, implying $0.28 and $0.56, but no effective date has been published. Our DeepSeek API pricing hub carries whatever is live.

Put $0.28 against the numbers we publish on LLM pricing. Median output price across the 135 models we track is $3.60. Cheapest model inside our overall top 10 is Grok 4.5 at $6.00. DeepSeek is charging roughly a twentieth of what the cheapest frontier-ranked model charges to generate a token, and no benchmark argument below changes that.

Running it yourself became a real option rather than a licensing question. vLLM and SGLang each take one flag to switch DSpark speculative decoding on, and Unsloth's lossless quant lands at 162GB, needing roughly 168GB of memory to run. Hardware feasibility gets its own piece later this month; our self-host calculator covers the cost comparison until then.

The two Terminal-Benches are different tests

Version 2.1 is what DeepSeek reports. That 56.9 riding alongside 82.7 belongs to Terminal-Bench 2.0, measured on the April Preview checkpoint. Subtracting one from the other produces 25.8 and produces nothing else.

We hold those two lanes in separate registry fields precisely so a version bump cannot quietly become a capability gain. Toolathlon has the same problem this week: 70.3 is Toolathlon-Verified, while the older Toolathlon key on the same row reads 47.8. Different test, different number, no arithmetic between them.

One independent check exists on the like-for-like figure. Artificial Analysis measured the then-current deepseek-v4-flash endpoint at 61.8 on Terminal-Bench 2.1 in its July 27 snapshot, four days before 0731 shipped. Land that against DeepSeek's own Preview column and it agrees to the decimal.

Two parties, same checkpoint, same benchmark version, same answer. That is the strongest evidence this launch has produced, and it argues for the smaller gain.

DeepSeek's own table does not claim parity with Opus 4.8

Alongside the version confusion runs a second claim: V4 Flash matches Claude Opus 4.8 at a fraction of the price. DeepSeek published a nine-row comparison against Opus 4.8. Opus 4.8 leads all nine.

Table 1
Benchmark V4 Flash 0731 V4 Flash Preview V4 Pro Preview GLM-5.2 Opus 4.8
Terminal-Bench 2.1 82.7 61.8 72.1 81.0 85.0
NL2Repo 54.2 39.4 38.5 48.9 69.7
Cybergym 76.7 38.7 52.7 83.1
DeepSWE 54.4 7.3 12.8 46.2 58.0
Toolathlon-Verified 70.3 49.7 55.9 59.9 76.2
Agents' Last Exam 25.2 15.8 16.5 23.8 25.7
AutomationBench Public 25.1 10.8 12.8 12.9 27.2
DSBench-FullStack † 68.7 37.0 41.8 61.8 71.6
DSBench-Hard † 59.6 25.8 31.1 54.5 71.7

† DSBench-FullStack and DSBench-Hard are DeepSeek's internal test sets. Source: DeepSeek-V4-Flash-0731 model card, retrieved August 1, 2026.

"Broadly competitive with the strongest proprietary models available" is how DeepSeek words it. Fair description of that table. Parity is not.

Width of the gaps should drive any porting decision. NL2Repo sits 15.5 points back and DSBench-Hard 12.1 points back, both repository-scale coding work. Agents' Last Exam finishes within half a point. Close on one agentic suite and far behind on repository tasks describes a routing rule, not a replacement.

Two of our own rows were wrong

Sourcing this piece turned up two errors in our data. Both are corrected as of August 1.

Those three 0731 Flash rows carried Proprietary, from a July 31 check that found no public weights. Re-checking a day later found them on Hugging Face under MIT. Flipping the rows moves our tracked open-weight share from 149 of 297 models to 152 of 297, and you can read the current figure on our open-source LLM statistics page.

Subtler was the second. Artificial Analysis's 61.8 sat on the Max row in the same Terminal-Bench 2.1 lane as DeepSeek's 82.7, for one model. Since AA took that measurement before 0731 existed, it describes a superseded checkpoint. We removed it instead of publishing a row asserting two Terminal-Bench 2.1 scores.

Concretely, the Max row still reports 86.2 on MMLU-Pro, 88.1 on GPQA Diamond, and 79 on SWE-bench Verified. Each of those is an April Preview measurement from Table 7 of the DeepSeek-V4 technical report, carried forward because DeepSeek did not rerun those lanes. Set them beside a Terminal-Bench score the lab did rerun and one row is describing two checkpoints.

Neither fix promotes V4 Flash into our overall ranking, and that is deliberate. General lanes on these rows still carry April Preview values, because DeepSeek did not rerun them for 0731, so they stay marked ranking-ineligible — visible, dated, and outside a composite they would distort. Quote a V4 Flash position in an overall ranking today and you are quoting a mixed snapshot.

An independent run decides this

Two of the nine benchmarks above are DeepSeek's internal sets, so those rows cannot be reproduced outside the lab.

Worse for verification, DeepSeek evaluated the Code Agent rows with "the minimal mode of DeepSeek Harness (to be released)." Nobody has that harness. Until it lands, every number in the table is provider-run and unreplicable, including the 82.7 this article defends.

So watch for an independent Terminal-Bench 2.1 run on 0731. Artificial Analysis has Preview at 61.8, DeepSeek claims 82.7 for the successor, and the next AA refresh becomes the first outside test of a 20.9-point gain currently resting on software nobody else can run.

Land it near 82.7 and this $0.28-output open-weight checkpoint got twenty points better in three months. Land it near 70 and the story was a harness.

Questions we got

Did DeepSeek V4 Flash gain 25.8 points on Terminal-Bench?

No. DeepSeek's model card reports 82.7 for the 0731 release and 61.8 for the April Preview on Terminal-Bench 2.1, a gain of 20.9 points. The 56.9 figure circulating alongside 82.7 is the Preview's Terminal-Bench 2.0 result, so the widely shared 25.8-point jump compares two different benchmark versions.

Is DeepSeek V4 Flash 0731 open weight?

Yes. DeepSeek published the 0731 weights on Hugging Face under the MIT License, and our registry moved the three Flash rows to Open Weight on August 1 after re-checking. The API serves the same checkpoint at $0.14 input and $0.28 output per million tokens, so both paths are available.

Does DeepSeek V4 Flash match Claude Opus 4.8?

Not on DeepSeek's own numbers. The model card compares the two across nine benchmarks and Opus 4.8 leads every row, from 85.0 against 82.7 on Terminal-Bench 2.1 to 71.7 against 59.6 on DSBench-Hard. DeepSeek describes the model as broadly competitive with strong proprietary systems, not as an equal.

How much does DeepSeek V4 Flash cost?

The regular rate is $0.14 per million input tokens, $0.0028 for cache hits, and $0.28 per million output tokens. DeepSeek has announced that every billing item will cost double during two Beijing-time windows each day, which would imply $0.28 and $0.56, but it has not published an effective date.

Why is DeepSeek V4 Flash not in the overall ranking?

Its rows are marked ranking-ineligible because the general benchmark lanes still carry April Preview values that DeepSeek did not rerun for 0731. Only the launch-day Code Agent table is fresh, and it was produced with an unreleased harness. The rows stay visible and dated rather than entering a composite they would distort.

Reader questions

Frequently asked questions

01Did DeepSeek V4 Flash gain 25.8 points on Terminal-Bench?

No. DeepSeek's model card reports 82.7 for the 0731 release and 61.8 for the April Preview on Terminal-Bench 2.1, a gain of 20.9 points. The 56.9 figure circulating alongside 82.7 is the Preview's Terminal-Bench 2.0 result, so the widely shared 25.8-point jump compares two different benchmark versions.

02Is DeepSeek V4 Flash 0731 open weight?

Yes. DeepSeek published the 0731 weights on Hugging Face under the MIT License, and our registry moved the three Flash rows to Open Weight on August 1 after re-checking. The API serves the same checkpoint at $0.14 input and $0.28 output per million tokens, so both paths are available.

03Does DeepSeek V4 Flash match Claude Opus 4.8?

Not on DeepSeek's own numbers. The model card compares the two across nine benchmarks and Opus 4.8 leads every row, from 85.0 against 82.7 on Terminal-Bench 2.1 to 71.7 against 59.6 on DSBench-Hard. DeepSeek describes the model as broadly competitive with strong proprietary systems, not as an equal.

04How much does DeepSeek V4 Flash cost?

The regular rate is $0.14 per million input tokens, $0.0028 for cache hits, and $0.28 per million output tokens. DeepSeek has announced that every billing item will cost double during two Beijing-time windows each day, which would imply $0.28 and $0.56, but it has not published an effective date.

05Why is DeepSeek V4 Flash not in the overall ranking?

Its rows are marked ranking-ineligible because the general benchmark lanes still carry April Preview values that DeepSeek did not rerun for 0731. Only the launch-day Code Agent table is fresh, and it was produced with an unreleased harness. The rows stay visible and dated rather than entering a composite they would distort.

Source ledger

External sources linked in this article

4
  1. 01Weights went to Hugging Face
  2. 02Pricing did not move
  3. 03Unsloth's lossless quant
  4. 04Artificial Analysis

Share or save

Share on XShare on LinkedIn

Keep reading

All research

Choose the right model before an expensive mistake. One weekly recommendation: what to choose, what costs less, and what is not worth switching to.

Read a sample issue

Join 2,000+ readers.