React Native Evals
We show this table for reference; we do not rank on it.
An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.
Benchmark score on React Native Evals — October 6, 2026
We compile the React Native Evals rows from benchmark-owner or independent runs. Composer 2 leads the table at 96.1%, followed by Composer 2 Fast (94.9%) and GPT-5.4 (85.3%). We do not use these results to rank models overall.
Composer 2
Cursor
Composer 2 Fast
Cursor
GPT-5.4
OpenAI
16 modelsCodingCurrentDisplay onlyUpdated October 6, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Composer 2Cursor | 96.1% | Not reported | Closed |
| 2 | Composer 2 FastCursor | 94.9% | Not reported | Closed |
| 3 | GPT-5.4OpenAI | 85.3% | Not reported | Closed |
| 4 | GPT-5.5OpenAI | 84.7% | Not reported | Closed |
| 5 | Claude Opus 4.6Anthropic | 84.1% | Not reported | Closed |
| 6 | Claude Opus 4.7Anthropic | 82.8% | Not reported | Closed |
| 7 | Claude Sonnet 4.6Anthropic | 80.6% | Not reported | Closed |
| 8 | Gemini 3.1 ProGoogle | 78.9% | Not reported | Closed |
| 9 | Kimi K2.5Moonshot AI | 77.2% | Not reported | Open |
| 10 | Gemma 4 31BGoogle | 75.2% | Not reported | Open |
| 11 | GLM-5Z.AI | 74.8% | Not reported | Open |
| 12 | Grok 4xAI | 72.6% | Not reported | Closed |
| 13 | GPT-OSS 120BOpenAI | 71.6% | Not reported | Open |
| 14 | DeepSeek V3.2DeepSeek | 71.5% | Not reported | Open |
| 15 | MiniMax M2.7MiniMax | 71.4% | Not reported | Open |
| 16 | GPT-OSS 20BOpenAI | 71% | Not reported | Open |
Among the reported React Native Evals rows, Composer 2 is first at 96.1%. The third row is 10.8 points behind. The broader top-10 range is 20.9 points, so the table still separates the published systems.
16 models have been evaluated on React Native Evals. The benchmark falls in the Coding category. React Native Evals is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About React Native Evals
Year
2026
Tasks
React Native app implementation tasks
Format
Framework-specific app development evaluation
Difficulty
Production mobile app engineering
React Native Evals focuses on framework-specific mobile work that generic coding benchmarks often miss. The public dashboard groups tasks into areas like navigation, animation, and async state, with repeated runs and cost tracking across models.
Freshness and provenance
Version
React Native Evals 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does React Native Evals measure?
An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.
Which model scores highest on React Native Evals?
Composer 2 by Cursor currently leads with a score of 96.1% on React Native Evals.
How many models are evaluated on React Native Evals?
16 AI models have published results on React Native Evals in the BenchLM catalog.
Compare top models on React Native Evals
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.