# Best LLMs for Coding — September 2026 Leaderboard

> As of September 2026, Claude Fable 5.1 leads BenchLM's coding leaderboard with a weighted score of 83.9.

- **Last verified:** September 15, 2026
- Canonical page: https://benchlm.ai/coding
- **Ranking coverage:** 152 category-ranked models from 486 tracked models
- **Category weight:** 20% of the overall BenchLM score

## Current ranking

| Rank | Model | Creator | Weighted score | Published category rows | Exact-source rows (all categories) |
|------|-------|---------|----------------|----------------|-------------------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 83.9 | 12 | 33 total |
| 2 | [Claude Fable 5](/models/claude-fable) | Anthropic | 77 | 12 | 31 total |
| 3 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 75.7 | 19 | 84 total |
| 4 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 74.5 | 5 | 34 total |
| 5 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 74.4 | 14 | 47 total |
| 6 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 67.8 | 15 | 52 total |
| 7 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 67.6 | 11 | 44 total |
| 8 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 67 | 11 | 44 total |
| 9 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 66.8 | 10 | 38 total |
| 10 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 66.5 | 9 | 29 total |
| 11 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 66.4 | 12 | 44 total |
| 12 | [Grok 4.6](/models/grok-4-6) | xAI | 66.2 | 11 | 23 total |
| 13 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 64 | 13 | 32 total |
| 14 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 62.8 | 5 | 11 total |
| 15 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 62.7 | 8 | 30 total |
| 16 | [Grok 4.5](/models/grok-4-5) | xAI | 62.1 | 9 | 24 total |
| 17 | [GPT-5.3 Codex](/models/gpt-5-3-codex) | OpenAI | 62 | 4 | 14 total |
| 18 | [GLM-5.3](/models/glm-5-3) | Z.AI | 61.4 | 15 | 32 total |
| 19 | [GLM-5.2](/models/glm-5-2) | Z.AI | 60.9 | 10 | 30 total |
| 20 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 60.8 | 13 | 55 total |
| 21 | [Muse Spark 1.2](/models/muse-spark-1-2) | Meta | 60.1 | 7 | 16 total |
| 22 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 59.9 | 2 | 18 total |
| 23 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 59.7 | 6 | 32 total |
| 24 | [Hy4 preview](/models/hy4-preview) | Tencent | 59.2 | 8 | 24 total |
| 25 | [Muse Spark](/models/muse-spark) | Meta | 58.9 | 5 | 30 total |
| 26 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 58.7 | 7 | 18 total |
| 27 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Google | 58.5 | 6 | 16 total |
| 28 | [Hy3](/models/hy3) | Tencent | 58.1 | 2 | 7 total |
| 29 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 58.1 | 7 | 25 total |
| 30 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 58 | 7 | 29 total |
| 31 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 57.7 | 4 | 20 total |
| 32 | [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) | Alibaba | 57.6 | 4 | 10 total |
| 33 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 57.4 | 10 | 38 total |
| 34 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 57.2 | 7 | 30 total |
| 35 | [Claude Opus 4.6 (Adaptive)](/models/claude-opus-4-6-thinking) | Anthropic | 56.8 | 1 | 1 total |
| 36 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 56.7 | 5 | 49 total |
| 37 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 56.5 | 8 | 37 total |
| 38 | [GLM-5.1](/models/glm-5-1) | Z.AI | 56.3 | 9 | 31 total |
| 39 | [GLM-5](/models/glm-5) | Z.AI | 56.1 | 6 | 40 total |
| 40 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | Xiaomi | 55.4 | 6 | 20 total |
| 41 | [MiMo-V2-Pro](/models/mimo-v2-pro) | Xiaomi | 55.3 | 1 | 7 total |
| 42 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 53.9 | 5 | 40 total |
| 43 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 53.9 | 10 | 42 total |
| 44 | [SWE-1.7](/models/swe-1-7) | Cognition | 53.2 | 3 | 4 total |
| 45 | [GLM-5V-Turbo](/models/glm-5v-turbo) | Z.AI | 52.8 | 0 | 8 total |
| 46 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 52.7 | 9 | 61 total |
| 47 | [Apodex 1.1](/models/apodex-1-1) | Apodex | 52.5 | 4 | 18 total |
| 48 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 52.3 | 8 | 30 total |
| 49 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 52.3 | 17 | 36 total |
| 50 | [Atria Dawn Preview](/models/atria-dawn-preview) | Shanghai Artificial Intelligence Laboratory | 51.6 | 2 | 12 total |
| 51 | [GPT-5.2-Codex](/models/gpt-5-2-codex) | OpenAI | 51.6 | 1 | 9 total |
| 52 | [Kimi K2.7 Code](/models/kimi-k2-7-code) | Moonshot AI | 50.9 | 7 | 11 total |
| 53 | [Kimi K2.5 (Reasoning)](/models/kimi-k2-5-reasoning) | Moonshot AI | 50.8 | 3 | 8 total |
| 54 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 50.8 | 12 | 34 total |
| 55 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 50.7 | 1 | 9 total |
| 56 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 50.6 | 9 | 48 total |
| 57 | [Composer 2.5](/models/composer-2-5) | Cursor | 50.5 | 5 | 5 total |
| 58 | [MAI-Thinking-1](/models/mai-thinking-1) | Microsoft | 50.5 | 4 | 13 total |
| 59 | [Ornith-1.0-397B](/models/ornith-1-0-397b) | DeepReinforce AI | 50.5 | 5 | 7 total |
| 60 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 50.4 | 3 | 35 total |
| 61 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 50 | 2 | 20 total |
| 62 | [BTL-4](/models/btl-4) | Bad Theory Labs | 50 | 2 | 2 total |
| 63 | [DeepSeek V3.2](/models/deepseek-v3-2) | DeepSeek | 50 | 2 | 12 total |
| 64 | [MiMo-V2-Flash](/models/mimo-v2-flash) | Xiaomi | 49.5 | 2 | 8 total |
| 65 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 49.4 | 12 | 47 total |
| 66 | [Composer 2](/models/composer-2) | Cursor | 49 | 4 | 5 total |
| 67 | [MiniMax M3](/models/minimax-m3) | MiniMax | 48.7 | 12 | 32 total |
| 68 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 48.6 | 13 | 30 total |
| 69 | [MiniMax M2.5](/models/minimax-m2-5) | MiniMax | 48.5 | 1 | 1 total |
| 70 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 48.3 | 6 | 23 total |
| 71 | [GPT-5 (medium)](/models/gpt-5-medium) | OpenAI | 48.3 | 0 | 7 total |
| 72 | [DeepSeek V3.2 (Thinking)](/models/deepseek-v3-2-thinking) | DeepSeek | 47.8 | 1 | 1 total |
| 73 | [GLM-4.7](/models/glm-4-7) | Z.AI | 47.7 | 5 | 16 total |
| 74 | [Step 3.7 Flash](/models/step-3-7-flash) | StepFun | 47.3 | 4 | 19 total |
| 75 | [Mercury 2.5](/models/mercury-2-5) | Inception | 47.2 | 1 | 7 total |
| 76 | [GLM-4.6](/models/glm-4-6) | Z.AI | 47.1 | 1 | 9 total |
| 77 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 46.9 | 8 | 52 total |
| 78 | [ZAYA1-74B-Preview](/models/zaya1-74b-preview) | Zyphra | 46.8 | 2 | 6 total |
| 79 | [Qwen3.5 Flash](/models/qwen3-5-flash) | Alibaba | 46.8 | 0 | 2 total |
| 80 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 46.8 | 7 | 18 total |
| 81 | [Ornith-1.0-35B](/models/ornith-1-0-35b) | DeepReinforce AI | 46.7 | 5 | 7 total |
| 82 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 46.4 | 3 | 20 total |
| 83 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 46.4 | 5 | 16 total |
| 84 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 46.2 | 7 | 25 total |
| 85 | [Ornith-1.0-9B](/models/ornith-1-0-9b) | DeepReinforce AI | 46.1 | 5 | 7 total |
| 86 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | xAI | 46.1 | 1 | 8 total |
| 87 | [Laguna S 2.1](/models/laguna-s-2-1) | Poolside | 46 | 4 | 6 total |
| 88 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 45.6 | 2 | 20 total |
| 89 | [o3-mini](/models/o3-mini) | OpenAI | 45.3 | 1 | 7 total |
| 90 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 45 | 8 | 49 total |
| 91 | [GPT-5.1](/models/gpt-5-1) | OpenAI | 45 | 2 | 11 total |
| 92 | [GPT-5.1-Codex](/models/gpt-5-1-codex) | OpenAI | 44.9 | 1 | 9 total |
| 93 | [K-Exaone](/models/k-exaone) | LG AI Research | 44.5 | 1 | 7 total |
| 94 | [Hy3 Preview](/models/hy3-preview) | Tencent | 44.4 | 5 | 12 total |
| 95 | [o1](/models/o1) | OpenAI | 44.4 | 1 | 10 total |
| 96 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 44.2 | 2 | 19 total |
| 97 | [o1-preview](/models/o1-preview) | OpenAI | 44.1 | 1 | 1 total |
| 98 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 43.8 | 8 | 32 total |
| 99 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 43.7 | 4 | 10 total |
| 100 | [Command A+](/models/command-a-plus) | Cohere | 43.6 | 2 | 12 total |
| 101 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 43.5 | 4 | 13 total |
| 102 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 43.3 | 6 | 19 total |
| 103 | [Claude 4.1 Opus](/models/claude-4-1-opus) | Anthropic | 42.9 | 1 | 2 total |
| 104 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 42.8 | 6 | 26 total |
| 105 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 42.6 | 3 | 22 total |
| 106 | [Gemma 4 31B](/models/gemma-4-31b) | Google | 42.4 | 4 | 15 total |
| 107 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 41.8 | 8 | 46 total |
| 108 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 41.4 | 3 | 17 total |
| 109 | [GPT-4.1](/models/gpt-4-1) | OpenAI | 41.3 | 1 | 13 total |
| 110 | [Gemma 4 26B A4B](/models/gemma-4-26b-a4b) | Google | 41.3 | 1 | 12 total |
| 111 | [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) | xAI | 41.2 | 1 | 8 total |
| 112 | [Inkling](/models/inkling) | Thinking Machines Lab | 41.1 | 8 | 31 total |
| 113 | [Mistral Small 4](/models/mistral-small-4) | Mistral | 41 | 2 | 9 total |
| 114 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 40.6 | 9 | 31 total |
| 115 | [Ling 2.6 Flash](/models/ling-2-6-flash) | InclusionAI | 40.5 | 2 | 10 total |
| 116 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 40.4 | 3 | 22 total |
| 117 | [GPT-5 mini](/models/gpt-5-mini) | OpenAI | 40.3 | 1 | 1 total |
| 118 | [Claude 4 Sonnet](/models/claude-4-sonnet) | Anthropic | 40.1 | 1 | 9 total |
| 119 | [DeepSeek V3](/models/deepseek-v3) | DeepSeek | 39.5 | 4 | 9 total |
| 120 | [Gemma 4 E4B](/models/gemma-4-e4b) | Google | 39.4 | 1 | 8 total |
| 121 | [Mistral Medium 3.5 128B](/models/mistral-medium-3-5-128b) | Mistral | 39.2 | 4 | 17 total |
| 122 | [Grok Build 0.1](/models/grok-build-0-1) | xAI | 39.1 | 0 | 0 total |
| 123 | [LFM2.5-2.6B](/models/lfm2-5-2-6b) | LiquidAI | 39 | 3 | 13 total |
| 124 | [Gemma 4 E2B](/models/gemma-4-e2b) | Google | 38.8 | 1 | 8 total |
| 125 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 38.7 | 3 | 9 total |
| 126 | [Gemini 3.1 Flash-Lite](/models/gemini-3-1-flash-lite) | Google | 38.2 | 3 | 7 total |
| 127 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 37.2 | 5 | 25 total |
| 128 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 37.2 | 8 | 20 total |
| 129 | [GPT-4.1 mini](/models/gpt-4-1-mini) | OpenAI | 35.6 | 2 | 13 total |
| 130 | [Gemini 1.5 Pro](/models/gemini-1-5-pro) | Google | 34.2 | 1 | 3 total |
| 131 | [Claude 3 Opus](/models/claude-3-opus) | Anthropic | 34 | 1 | 2 total |
| 132 | [Grok 4.3](/models/grok-4-3) | xAI | 33.7 | 5 | 20 total |
| 133 | [Gemma 3 27B](/models/gemma-3-27b) | Google | 33.1 | 2 | 9 total |
| 134 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 31.8 | 4 | 14 total |
| 135 | [GPT-4o mini](/models/gpt-4o-mini) | OpenAI | 31.7 | 1 | 5 total |
| 136 | [Llama 4 Scout](/models/llama-4-scout) | Meta | 31.1 | 2 | 10 total |
| 137 | [GPT-4.1 nano](/models/gpt-4-1-nano) | OpenAI | 31.1 | 1 | 12 total |
| 138 | [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) | NVIDIA | 30.3 | 6 | 14 total |
| 139 | [Ministral 3 14B](/models/ministral-3-14b) | Mistral | 30.2 | 0 | 0 total |
| 140 | [Grok Code Fast 1](/models/grok-code-fast-1) | xAI | 30 | 1 | 6 total |
| 141 | [Laguna M.1](/models/laguna-m-1) | Poolside | 29.7 | 6 | 10 total |
| 142 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 28.4 | 6 | 22 total |
| 143 | [GPT-4 Turbo](/models/gpt-4-turbo) | OpenAI | 27.8 | 1 | 1 total |
| 144 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 27.4 | 4 | 10 total |
| 145 | [Trinity-Large-Thinking](/models/trinity-large-thinking) | Arcee AI | 27.4 | 3 | 12 total |
| 146 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 26.6 | 9 | 26 total |
| 147 | [Llama 4 Maverick](/models/llama-4-maverick) | Meta | 25.7 | 2 | 10 total |
| 148 | [Mistral Large 3](/models/mistral-large-3) | Mistral | 25.4 | 2 | 9 total |
| 149 | [Laguna XS.2](/models/laguna-xs-2) | Poolside | 24.7 | 6 | 10 total |
| 150 | [Mercury 2](/models/mercury-2) | Inception | 24.6 | 0 | 0 total |
| 151 | [Ministral 3 8B](/models/ministral-3-8b) | Mistral | 21.1 | 0 | 0 total |
| 152 | [Ministral 3 3B](/models/ministral-3-3b) | Mistral | 19 | 0 | 0 total |

## Decision-ready shortlist

- #1 [Claude Fable 5.1](/models/claude-fable-5-1) — 83.9 weighted score, Proprietary, 1M context.
- #2 [Claude Fable 5](/models/claude-fable) — 77 weighted score, Proprietary, 1M+ context.
- #3 [Claude Opus 5](/models/claude-opus-5) — 75.7 weighted score, Proprietary, null context.
- #4 [GPT-6 Astra](/models/gpt-6-astra) — 74.5 weighted score, Proprietary, 1.05M context.
- #5 [GPT-5.6 Sol](/models/gpt-5-6-sol) — 74.4 weighted score, Proprietary, 1.05M context.

## Benchmarks in this category

### [HumanEval](/benchmarks/humaneval) (Evaluating Large Language Models Trained on Code)

A set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check, but BenchLM's current exact-source table is too small to support a broad frontier-coding verdict.

- Ranking status: Display only
- Year: 2021
- Format: Python function generation
- Difficulty: Introductory to intermediate programming

### [SWE-bench Verified](/benchmarks/swe-bench-verified) (Software Engineering Benchmark Verified)

A curated, human-verified subset of SWE-bench that tests models on resolving real GitHub issues from popular open-source Python repositories like Django, Flask, and scikit-learn.

- Ranking status: Weighted (10% of this category)
- Year: 2024
- Format: Code patch generation
- Difficulty: Professional software engineering

### [LiveCodeBench](/benchmarks/livecodebench) (LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code)

A continuously updated coding benchmark built from newly collected LeetCode, AtCoder, and Codeforces problems. Fresh problem windows reduce one contamination path, but results still need a release and setup check.

- Ranking status: Weighted (15% of this category)
- Year: 2024
- Format: Competitive-programming evaluation
- Difficulty: Competitive programming level

### [LiveCodeBench Pro](/benchmarks/livecodebench-pro) (LiveCodeBench Pro)

A harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with quarter-specific public leaderboards and difficulty-aware reporting.

- Ranking status: Display only
- Year: 2025
- Format: Competitive programming
- Difficulty: High-end contest programming

### [FLTEval](/benchmarks/flteval) (FLTEval)

A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.

- Ranking status: Display only
- Year: 2026
- Format: Lean 4 repository task completion
- Difficulty: Formal verification / proof engineering

### [SWE-bench Pro](/benchmarks/swe-bench-pro) (SWE-bench Pro)

A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.

- Ranking status: Weighted (25% of this category)
- Year: 2025
- Format: Repository task completion
- Difficulty: Long-horizon professional engineering

### [SWE-Rebench](/benchmarks/swe-rebench) (SWE-Rebench)

A continuously updated software engineering benchmark by Nebius using fresh GitHub issues to avoid contamination. Models are evaluated 5 times per problem under a fixed ReAct scaffolding; the Resolved Rate (best pass@1) is reported.

- Ranking status: Weighted (10% of this category)
- Year: 2026
- Format: Code patch generation
- Difficulty: Professional software engineering

### [SWE Multilingual](/benchmarks/swe-bench-multilingual) (SWE Multilingual)

A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.

- Ranking status: Weighted (5% of this category)
- Year: 2026
- Format: Repository task completion
- Difficulty: Professional software engineering

### [Multi-SWE Bench](/benchmarks/multiswebench) (Multi-SWE Bench)

A multi-language software-engineering benchmark that measures repository-level bug fixing and implementation across more than one programming ecosystem.

- Ranking status: Display only
- Year: 2026
- Format: Repository task completion
- Difficulty: Professional software engineering

### [VIBE-Pro](/benchmarks/vibepro) (VIBE-Pro)

A repo-level code generation and full-project delivery benchmark spanning web, mobile, and simulation-style implementation tasks.

- Ranking status: Display only
- Year: 2026
- Format: Repository-level implementation benchmark
- Difficulty: End-to-end software delivery

### [NL2Repo](/benchmarks/nl2repo) (NL2Repo)

A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.

- Ranking status: Display only
- Year: 2026
- Format: Repository understanding benchmark
- Difficulty: System-level software comprehension

### [Vibe Code Bench](/benchmarks/vibecodebench) (Vibe Code Bench v1.1)

Vals.ai benchmark for evaluating whether models can build complete web applications from natural language specifications in a production-like development environment.

- Ranking status: Display only
- Year: 2026
- Format: Full-stack app implementation benchmark
- Difficulty: End-to-end software delivery

### [React Native Evals](/benchmarks/reactnativeevals) (React Native Evals)

An open benchmark for AI coding agents on real-world React Native implementation tasks, emphasizing working app behavior, recommended architecture choices, and strict constraint adherence.

- Ranking status: Display only
- Year: 2026
- Format: Framework-specific app development evaluation
- Difficulty: Production mobile app engineering

### [SWE-bench Verified*](/benchmarks/sweverifiedarcee) (SWE-bench Verified (mini-swe-agent-v2))

A display-only SWE-bench Verified reference from Arcee AI's Trinity-Large-Thinking comparison chart.

- Ranking status: Display only
- Year: 2026
- Format: Agent scaffold benchmark
- Difficulty: Professional software engineering

### [Spider 2.0-Lite](/benchmarks/spider2lite) (Spider 2.0-Lite)

A text-to-SQL benchmark over realistic warehouse-scale schemas, reported by Interfaze for model comparison.

- Ranking status: Display only
- Year: 2024
- Format: Execution accuracy
- Difficulty: Enterprise text-to-SQL

### [Bug Hunt Bench](/benchmarks/bug-hunt-bench) (Bug Hunt Bench)

A blind-graded coding-agent benchmark with 105 planted bugs across two production TypeScript repositories.

- Ranking status: Display only
- Year: 2026
- Format: Strict planted bugs fixed
- Difficulty: Blind production-repository bug finding and repair
