# React Native Evals

> An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.

Canonical page: https://benchlm.ai/benchmarks/reactnativeevals

- Category: [Coding](/coding)
- Last updated: September 15, 2026

## About React Native Evals

- Year: 2026
- Tasks: React Native app implementation tasks
- Format: Framework-specific app development evaluation
- Difficulty: Production mobile app engineering
- Paper: [React Native Evals](https://rn-evals.vercel.app/)

React Native Evals focuses on framework-specific mobile work that generic coding benchmarks often miss. The public dashboard groups tasks into areas like navigation, animation, and async state, with repeated runs and cost tracking across models.

React Native Evals is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (16 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Composer 2](/models/composer-2) | Cursor | 96.1% |
| 2 | [Composer 2 Fast](/models/composer-2-fast) | Cursor | 94.9% |
| 3 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 85.3% |
| 4 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 84.7% |
| 5 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 84.1% |
| 6 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 82.8% |
| 7 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 80.6% |
| 8 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 78.9% |
| 9 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 77.2% |
| 10 | [Gemma 4 31B](/models/gemma-4-31b) | Google | 75.2% |
| 11 | [GLM-5](/models/glm-5) | Z.AI | 74.8% |
| 12 | [Grok 4](/models/grok-4) | xAI | 72.6% |
| 13 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 71.6% |
| 14 | [DeepSeek V3.2](/models/deepseek-v3-2) | DeepSeek | 71.5% |
| 15 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 71.4% |
| 16 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 71% |

## FAQ

### What does React Native Evals measure?

An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.

### Which model scores highest on React Native Evals?

Composer 2 by Cursor currently leads with a score of 96.1% on React Native Evals.

### How many models are evaluated on React Native Evals?

16 AI models have been evaluated on React Native Evals on BenchLM.

### Does React Native Evals affect BenchLM's overall score?

Not directly. React Native Evals is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on React Native Evals

- [Composer 2 vs Composer 2 Fast](/compare/composer-2-vs-composer-2-fast)
- [Composer 2 Fast vs GPT-5.4](/compare/composer-2-fast-vs-gpt-5-4)
- [GPT-5.4 vs GPT-5.5](/compare/gpt-5-4-vs-gpt-5-5)
- [GPT-5.5 vs Claude Opus 4.6](/compare/claude-opus-4-6-vs-gpt-5-5)

## Related Reading

- [React Native Evals benchmark explainer](/blog/posts/react-native-evals-mobile-benchmark)
