# AI Agent Evaluations for Next.js (Next.js Evals)

> A Vercel benchmark for AI coding agents on Next.js code generation and migration tasks, reporting success rate, average execution time, and an AGENTS.md documentation-assisted split.

Canonical page: https://benchlm.ai/benchmarks/nextjsevals

- Category: [Coding](/coding)
- Last updated: June 1, 2026

## About Next.js Evals

- Year: 2026
- Tasks: 24 Next.js code generation and migration tasks
- Format: Agent task completion with withheld Vitest assertions
- Difficulty: Framework-specific web application engineering
- Paper: [AI Agent Evaluations | Next.js](https://nextjs.org/evals)

Next.js Evals focuses on framework-specific web engineering tasks such as Pages Router to App Router migration, server actions, cache directives, proxy middleware, async cookies and headers, and other current Next.js patterns. BenchLM mirrors the public leaderboard as display-only because rows combine model choice with an agent harness.

Next.js Evals is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (18 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Composer 2.5](/models/composer-2-5) | Cursor | 92% |
| 2 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 88% |
| 3 | [GPT-5.5 Pro](/models/gpt-5-5-pro) | OpenAI | 83% |
| 4 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 83% |
| 5 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 83% |
| 6 | [MiniMax M3](/models/minimax-m3) | MiniMax | 75% |
| 7 | [GLM-5.1](/models/glm-5-1) | Z.AI | 75% |
| 8 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 75% |
| 9 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 75% |
| 10 | [Composer 2](/models/composer-2) | Cursor | 75% |
| 11 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 75% |
| 12 | [Gemini 3.0 Pro Preview](https://nextjs.org/evals) | Google | 67% |
| 13 | [Cursor Composer 1.5](https://nextjs.org/evals) | Cursor | 67% |
| 14 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 58% |
| 15 | [GPT-5.2-Codex](/models/gpt-5-2-codex) | OpenAI | 58% |
| 16 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 50% |
| 17 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 50% |
| 18 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 21% |

## FAQ

### What does Next.js Evals measure?

A Vercel benchmark for AI coding agents on Next.js code generation and migration tasks, reporting success rate, average execution time, and an AGENTS.md documentation-assisted split.

### Which model leads the published Next.js Evals snapshot?

Composer 2.5 currently leads the published Next.js Evals snapshot with a score of 92%.

### How many models are evaluated on Next.js Evals?

The June 1, 2026 contains 18 AI models.

### Does Next.js Evals affect BenchLM's overall score?

Not directly. Next.js Evals is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
