# OpenHands Index

> A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.

Canonical page: https://benchlm.ai/benchmarks/openhandsindex

- Category: [Agentic](/agentic)
- Last updated: September 18, 2026 snapshot

## About OpenHands Index

- Year: 2025
- Tasks: SWE-bench Verified, SWE-bench Multimodal, Commit0, SWT-bench Verified, and GAIA
- Format: Macro-average across five coding-agent categories
- Difficulty: Real-world software engineering agent tasks
- Paper: [OpenHands Index methodology](https://index.openhands.dev/about)

BenchLM mirrors the official OpenHands Index REST API as a display-only agentic software-engineering benchmark. The source reports average agent score, cost, runtime, per-category scores, logs, and visualizations for each model and SDK version.

OpenHands Index is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (33 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5](/models/claude-fable) | Anthropic | 81.0% |
| 2 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 71.9% |
| 3 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 69.7% |
| 4 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 66.7% |
| 5 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 65.9% |
| 6 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 64.3% |
| 7 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 62.6% |
| 8 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 60.6% |
| 9 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 60.6% |
| 10 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 58.8% |
| 11 | [GPT-5.2-Codex](/models/gpt-5-2-codex) | OpenAI | 58.3% |
| 12 | [GLM-5.1](/models/glm-5-1) | Z.AI | 58.2% |
| 13 | [MiniMax M3](/models/minimax-m3) | MiniMax | 57.2% |
| 14 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 57.1% |
| 15 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 53.0% |
| 16 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 52.9% |
| 17 | [GLM-5](/models/glm-5) | Z.AI | 49.4% |
| 18 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 49.2% |
| 19 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 49.0% |
| 20 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 49.0% |
| 21 | [DeepSeek V3.2 (Thinking)](/models/deepseek-v3-2-thinking) | DeepSeek | 45.7% |
| 22 | [MiniMax M2.5](/models/minimax-m2-5) | MiniMax | 45.2% |
| 23 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 44.5% |
| 24 | [Qwen3 Coder Next](https://index.openhands.dev/api/leaderboard/model/Qwen3-Coder-Next) | Alibaba | 43.8% |
| 25 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 43.4% |
| 26 | [GLM-4.7](/models/glm-4-7) | Z.AI | 42.3% |
| 27 | [MiniMax M2.1](https://index.openhands.dev/api/leaderboard/model/MiniMax-M2.1) | MiniMax | 41.2% |
| 28 | [Kimi K2.5 (Reasoning)](/models/kimi-k2-5-reasoning) | Moonshot AI | 41.0% |
| 29 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 40.7% |
| 30 | [Qwen3.5 Flash](/models/qwen3-5-flash) | Alibaba | 38.1% |
| 31 | [Nemotron 3 Super 120B A12B](/models/nemotron-3-super-120b-a12b) | NVIDIA | 36.2% |
| 32 | [Trinity-Large-Thinking](/models/trinity-large-thinking) | Arcee AI | 32.1% |
| 33 | [Qwen3 Coder 480B A35B](https://index.openhands.dev/api/leaderboard/model/Qwen3-Coder-480B) | Alibaba | 30.9% |

## FAQ

### What does OpenHands Index measure?

A holistic coding-agent benchmark that evaluates AI agents across issue resolution, frontend work, greenfield development, testing, and information gathering.

### Which model leads the published OpenHands Index snapshot?

Claude Fable 5 currently leads the published OpenHands Index snapshot with a score of 81.0%.

### How many models are evaluated on OpenHands Index?

The September 18, 2026 snapshot contains 33 AI models.

### Does OpenHands Index affect BenchLM's overall score?

Not directly. OpenHands Index is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
