App-Bench
We show this table for reference; we do not rank on it.
A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.
Published percentile score on App-Bench — September 23, 2026 snapshot
We mirror the published published percentile score view for App-Bench. Orchids leads the public snapshot at 76.80%, followed by Claude Opus 4.5 (67.50%) and v0 (64.90%). We do not use these results to rank models overall.
Orchids
Orchids
Claude Opus 4.5
Anthropic
v0
Vercel
10 modelsCodingCurrentDisplay onlyUpdated September 23, 2026 snapshot
Published percentile score table (10 models)
ScoreHow we show App-Bench
We mirror AfterQuery's App-Bench table captured on September 23, 2026 snapshot. It compares 10 app-building systems on 6 full-stack projects using binary functional rubrics and reports the best result from three one-shot attempts per task.
The rows mix hosted app builders with CLI and IDE coding assistants. We preserve those system names and underlying model labels, but keep the table display-only because a tool-level result is not a controlled base-model comparison.
Snapshot
The published App-Bench snapshot places Orchids first at 76.80%. The third row is 11.90 points behind. The broader top-10 range is 76.80 points, so the table still separates the published systems.
10 models have been evaluated on App-Bench. The benchmark falls in the Coding category. App-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About App-Bench
Year
2025
Tasks
6 full-stack app-building tasks
Format
Best-of-three one-shot feature completion
Difficulty
Production-style full-stack application generation
App-Bench compares five hosted app builders and five CLI or IDE coding assistants. Each system receives three one-shot attempts per task, the best attempt is graded against binary functional requirements, and two developers reconcile disputed grades. We keep these tool-level rows display-only rather than treating them as base-model scores.
Freshness and provenance
Version
App-Bench 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does App-Bench measure?
A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.
Which model leads the published App-Bench snapshot?
Orchids currently leads the published App-Bench snapshot with 76.80% published percentile score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on App-Bench?
The September 23, 2026 snapshot snapshot contains 10 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.