Vals Vibe Code Bench 1-100 (Vibe Code Bench 1-100)
We mirror this table; we do not rank on it.
Can a model extend one working web application across a long sequence of dependent requests?
Vibe Code 1-100 score on Vibe Code Bench 1-100 — September 16, 2026
We mirror the published vibe code 1-100 score view for Vibe Code Bench 1-100. Claude Opus 5 leads the public snapshot at 28.53%, followed by Claude Fable 5.1 (28.00%) and GPT-6 Astra (27.64%). We do not use these results to rank models overall.
Claude Opus 5
Anthropic
OpenHands
OpenHands
Claude Fable 5.1
Anthropic
OpenHands
OpenHands
GPT-6 Astra
OpenAI
OpenHands
OpenHands
19 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated September 16, 2026
Vibe Code 1-100 score table (19 models)
ScoreHow Vibe Code Bench 1-100 is shown here
BenchLM mirrors the public Vals AI Vibe Code Bench 1-100 leaderboard captured from https://www.vals.ai/benchmarks/vcb-1-100 and updated by Vals on September 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Vibe Code Bench 1-100 is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Snapshot
The published Vibe Code Bench 1-100 snapshot places Claude Opus 5 first at 28.53%. The third row is 0.89 points behind. The broader top-10 range is 11.00 points, so the table still separates the published systems.
19 models have been evaluated on Vibe Code Bench 1-100. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vibe Code Bench 1-100 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Vibe Code Bench 1-100
Year
2026
Tasks
Long chains of dependent web-application change requests
Format
Accuracy score
Difficulty
Long-horizon software delivery
Each run starts from a working web application and applies a long chain of dependent change requests, so an early mistake compounds through the rest of the sequence. That is a different question from the single-build Vibe Code Bench v1.1 board, which is why BenchLM keeps it on its own key. Private task set, external runner, display only.
Freshness and provenance
Version
Vibe Code Bench 1-100 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Vibe Code Bench 1-100 measure?
Can a model extend one working web application across a long sequence of dependent requests?
Which model leads the published Vibe Code Bench 1-100 snapshot?
Claude Opus 5 currently leads the published Vibe Code Bench 1-100 snapshot with 28.53% vibe code 1-100 score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Vibe Code Bench 1-100?
The September 16, 2026 snapshot contains 19 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.