LiveCodeBench Pass@1 with Chain-of-Thought (LiveCodeBench Pass@1-COT)
This lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps them separate from generic and version-specific LiveCodeBench rows.
Data verified 25 confirmed releases in the last 30 daysStart the free Radar BriefHow to read this leaderboard
Editorial review by Glevd · 2026-07-15
Compare these rows within the published DeepSeek table. Do not merge them with pass@1 values that use another prompt policy, release window, sampling scheme, or execution configuration.
Operator receipt: 6 sourced rows are currently displayable on this page; the leading published row is DeepSeek V4 Pro 0813 at 93.5%.
Honest limit: All current rows come from one provider report and one model family. They show within-family differences under the reported setup, not a broad cross-provider coding ranking or repository-level software-engineering ability.
Benchmark score on LiveCodeBench Pass@1-COT — August 29, 2026
We mirror the published score view for LiveCodeBench Pass@1-COT. DeepSeek V4 Pro 0813 leads the public snapshot at 93.5%, followed by DeepSeek V4 Flash 0731 (91.6%) and DeepSeek V4 Pro (High) (89.8%). We do not use these results to rank models overall.
DeepSeek V4 Pro 0813
DeepSeek
DeepSeek V4 Flash 0731
DeepSeek
DeepSeek V4 Pro (High)
DeepSeek
Benchmark score table (6 models)
ScoreThe published LiveCodeBench Pass@1-COT snapshot places DeepSeek V4 Pro 0813 first at 93.5%. The third row is 3.7 points behind. The broader top-10 range is 38.3 points, so the table still separates the published systems.
6 models have been evaluated on LiveCodeBench Pass@1-COT. The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. LiveCodeBench Pass@1-COT is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About LiveCodeBench Pass@1-COT
Year
2026
Tasks
DeepSeek-V4 report evaluation window
Format
Pass@1-COT competitive programming results
Difficulty
Competitive programming level
The source reports six reasoning-effort and model variants under a Pass@1-COT column. BenchLM preserves that published metric label and source family without treating the rows as a controlled comparison against generic LiveCodeBench results.
BenchLM freshness & provenance
Version
LiveCodeBench Pass@1-COT 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does LiveCodeBench Pass@1-COT measure?
This lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps them separate from generic and version-specific LiveCodeBench rows.
Which model scores highest on LiveCodeBench Pass@1-COT?
DeepSeek V4 Pro 0813 by DeepSeek currently leads with a score of 93.5% on LiveCodeBench Pass@1-COT.
How many models are evaluated on LiveCodeBench Pass@1-COT?
6 AI models have been evaluated on LiveCodeBench Pass@1-COT on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.