LiveCodeBench, Vals AI run (LiveCodeBench (Vals))
We show this table for reference; we do not rank on it.
Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks.
Vals LiveCodeBench mirror score on LiveCodeBench (Vals) — September 1, 2026
We mirror the published vals livecodebench mirror score view for LiveCodeBench (Vals). Claude Fable 5.1 leads the public snapshot at 90.5%, followed by Claude Fable 5 (89.8%) and Gemini 3.8 Flash (89.5%). We do not use these results to rank models overall.
Claude Fable 5.1
Anthropic
Claude Fable 5
Anthropic
Gemini 3.8 Flash
high reasoning
143 modelsCoding8% of Coding reference weightCurrentUpdated September 1, 2026
Vals LiveCodeBench mirror score table (143 models)
ScoreHow Vals LiveCodeBench mirror is shown here
BenchLM mirrors the public Vals AI Vals LiveCodeBench mirror leaderboard captured from https://www.vals.ai/benchmarks/lcb and updated by Vals on September 1, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Each row keeps the harness and settings the source published. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking.
Snapshot
The published LiveCodeBench (Vals) snapshot places Claude Fable 5.1 first at 90.5%. The third row is 1.0 points behind. The broader top-10 range is 2.7 points, so many of the published results sit in a relatively narrow band.
143 models have been evaluated on LiveCodeBench (Vals). The benchmark falls in the Coding category. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking. BenchAlign v5.7 gives LiveCodeBench (Vals) 8% of the Coding reference weight, so it moves the Coding leaderboard and the overall ranking. Reference weights are relative weights in the calibrated model, not fixed shares of a score.
About LiveCodeBench (Vals)
Year
2026
Tasks
Competitive programming problems (easy, medium, hard)
Format
Pass@1 accuracy
Difficulty
Frontier coding
BenchLM mirrors the Vals AI board on a dedicated key so a provider-run row on the canonical key is never overwritten. Vals publishes per-task accuracy with standard error, latency, and cost for every model it runs. Admitted as independent third-party evidence in methodology v5.5 (2026-09-04).
Freshness and provenance
Version
LiveCodeBench (Vals) 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does LiveCodeBench (Vals) measure?
Vals AI’s independent implementation of the LiveCodeBench code-generation benchmark, run under one fixed harness across the models it tracks.
Which model leads the published LiveCodeBench (Vals) snapshot?
Claude Fable 5.1 currently leads the published LiveCodeBench (Vals) snapshot with 90.5% vals livecodebench mirror score. BenchLM shows the published table for reference; the per-model scores it stores feed the ranking.
How many models are evaluated on LiveCodeBench (Vals)?
The September 1, 2026 snapshot contains 143 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.