LiveCodeBench, Vals AI run (LiveCodeBench (Vals))
Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks.
Data verified 31 confirmed releases in the last 30 daysSee provider release alertsHow BenchLM shows Vals LiveCodeBench mirror
BenchLM mirrors the public Vals AI Vals LiveCodeBench mirror leaderboard captured from https://www.vals.ai/benchmarks/lcb and updated by Vals on September 1, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Vals LiveCodeBench mirror is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Snapshot
Vals LiveCodeBench mirror score on LiveCodeBench (Vals) — September 1, 2026
We mirror the published vals livecodebench mirror score view for LiveCodeBench (Vals). Claude Fable 5.1 leads the public snapshot at 90.5%, followed by Claude Fable 5 (89.8%) and Gemini 3.8 Flash (89.5%). We do not use these results to rank models overall.
Claude Fable 5.1
Anthropic
Claude Fable 5
Anthropic
Gemini 3.8 Flash
high reasoning
143 modelsCoding15% of category scoreCurrentUpdated September 1, 2026
Vals LiveCodeBench mirror score table (143 models)
ScoreThe published LiveCodeBench (Vals) snapshot places Claude Fable 5.1 first at 90.5%. The third row is 1.0 points behind. The broader top-10 range is 2.7 points, so many of the published results sit in a relatively narrow band.
143 models have been evaluated on LiveCodeBench (Vals). The benchmark falls in the Coding category. This category carries a 20% weight in BenchLM.ai's overall scoring system. Within that category, LiveCodeBench (Vals) contributes 15% of the category score, so strong performance here directly affects a model's overall ranking.
About LiveCodeBench (Vals)
Year
2026
Tasks
Competitive programming problems (easy, medium, hard)
Format
Pass@1 accuracy
Difficulty
Frontier coding
BenchLM mirrors the Vals AI board on a dedicated key so a provider-run row on the canonical key is never overwritten. Vals publishes per-task accuracy with standard error, latency, and cost for every model it runs. Admitted as independent third-party evidence in methodology v5.5 (2026-09-04).
BenchLM freshness & provenance
Version
LiveCodeBench (Vals) 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does LiveCodeBench (Vals) measure?
Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks.
Which model leads the published LiveCodeBench (Vals) snapshot?
Claude Fable 5.1 currently leads the published LiveCodeBench (Vals) snapshot with 90.5% vals livecodebench mirror score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on LiveCodeBench (Vals)?
The September 1, 2026 snapshot contains 143 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.