# App-Bench

> A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.

Canonical page: https://benchlm.ai/benchmarks/appbench

- Category: [Coding](/coding)
- Last updated: September 23, 2026 snapshot

## About App-Bench

- Year: 2025
- Tasks: 6 full-stack app-building tasks
- Format: Best-of-three one-shot feature completion
- Difficulty: Production-style full-stack application generation
- Paper: [App-Bench](https://appbench.ai/)

App-Bench compares five hosted app builders and five CLI or IDE coding assistants. Each system receives three one-shot attempts per task, the best attempt is graded against binary functional requirements, and two developers reconcile disputed grades. We keep these tool-level rows display-only rather than treating them as base-model scores.

App-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (10 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Orchids](https://www.afterquery.com/leaderboard/app-bench) | Orchids | 76.80% |
| 2 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 67.50% |
| 3 | [v0](https://www.afterquery.com/leaderboard/app-bench) | Vercel | 64.90% |
| 4 | [Bolt](https://www.afterquery.com/leaderboard/app-bench) | StackBlitz | 53.60% |
| 5 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 50.30% |
| 6 | [GPT-5.1-Codex-Max](/models/gpt-5-1-codex-max) | OpenAI | 38.40% |
| 7 | [Replit](https://www.afterquery.com/leaderboard/app-bench) | Replit | 35.10% |
| 8 | [Cursor (Composer 1)](https://www.afterquery.com/leaderboard/app-bench) | Cursor | 27.80% |
| 9 | [Lovable](https://www.afterquery.com/leaderboard/app-bench) | Lovable | 25.83% |
| 10 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 0.00% |

## FAQ

### What does App-Bench measure?

A six-task full-stack web-app benchmark that measures how much required functionality an AI builder or coding assistant delivers from one prompt without human code edits.

### Which model leads the published App-Bench snapshot?

Orchids currently leads the published App-Bench snapshot with a score of 76.80%.

### How many models are evaluated on App-Bench?

The September 23, 2026 snapshot contains 10 AI models.

### Does App-Bench affect BenchLM's overall score?

Not directly. App-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
