# NL2Repo

> A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.

Canonical page: https://benchlm.ai/benchmarks/nl2repo

- Category: [Coding](/coding)
- Last updated: September 27, 2026

## About NL2Repo

- Year: 2026
- Tasks: Natural language to repository tasks
- Format: Repository understanding benchmark
- Difficulty: System-level software comprehension
- Paper: [MiniMax M2.7: Early Echoes of Self-Evolution](https://www.minimax.io/news/minimax-m27-en)

MiniMax cites NL2Repo as a system-level engineering benchmark that rewards deep understanding of complex repositories and their operational structure.

NL2Repo is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (29 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 65.4% |
| 2 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 61.5% |
| 3 | [Ornith-1.5-397B](/models/ornith-1-5-397b) | Ornith AI | 59.5% |
| 4 | [Hy4 preview](/models/hy4-preview) | Tencent | 58.9% |
| 5 | [GLM-5.3](/models/glm-5-3) | Z.AI | 58% |
| 6 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Z.AI | 56.3% |
| 7 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 55.9% |
| 8 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 54.2% |
| 9 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 49.8% |
| 10 | [GLM-5.2](/models/glm-5-2) | Z.AI | 48.9% |
| 11 | [Qwen3.8-Omni-Flash](/models/qwen3-8-omni-flash) | Alibaba | 48.9% |
| 12 | [Ornith-1.0-397B](/models/ornith-1-0-397b) | DeepReinforce AI | 48.2% |
| 13 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 48.1% |
| 14 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 47.2% |
| 15 | [Seed 2.1 Pro](/models/seed-2-1-pro) | ByteDance | 47% |
| 16 | [Ornith-1.5-35B-A3B](/models/ornith-1-5-35b-a3b) | Ornith AI | 46.2% |
| 17 | [Seed 2.1 Turbo](/models/seed-2-1-turbo) | ByteDance | 43.7% |
| 18 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 43.2% |
| 19 | [Qwen 3.6 Max (preview)](/models/qwen3-6-max-preview) | Alibaba | 42.9% |
| 20 | [GLM-5.1](/models/glm-5-1) | Z.AI | 42.7% |
| 21 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 42.3% |
| 22 | [MiniMax M3](/models/minimax-m3) | MiniMax | 42.1% |
| 23 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 41.1% |
| 24 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 39.8% |
| 25 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 36.2% |
| 26 | [Ornith-1.0-35B](/models/ornith-1-0-35b) | DeepReinforce AI | 34.6% |
| 27 | [Ornith-1.5-9B](/models/ornith-1-5-9b) | Ornith AI | 32.4% |
| 28 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 29.4% |
| 29 | [Ornith-1.0-9B](/models/ornith-1-0-9b) | DeepReinforce AI | 27.2% |

## FAQ

### What does NL2Repo measure?

A repository-understanding benchmark that measures whether models can map natural-language requests onto the right code locations and system changes.

### Which model scores highest on NL2Repo?

DeepSeek V4.1 Flash by DeepSeek currently leads with a score of 65.4% on NL2Repo.

### How many models are evaluated on NL2Repo?

29 AI models have been evaluated on NL2Repo on BenchLM.

### Does NL2Repo affect BenchLM's overall score?

Not directly. NL2Repo is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on NL2Repo

- [DeepSeek V4.1 Flash vs DeepSeek V4 Pro 0813](/compare/deepseek-v4-1-flash-vs-deepseek-v4-pro-0813)
- [DeepSeek V4 Pro 0813 vs Ornith-1.5-397B](/compare/deepseek-v4-pro-0813-vs-ornith-1-5-397b)
- [Ornith-1.5-397B vs Hy4 preview](/compare/hy4-preview-vs-ornith-1-5-397b)
- [Hy4 preview vs GLM-5.3](/compare/glm-5-3-vs-hy4-preview)
