# DeepPlanning

> A long-horizon planning benchmark that tests whether agents can optimize under explicit time, budget, and feasibility constraints.

Canonical page: https://benchlm.ai/benchmarks/deepplanning

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About DeepPlanning

- Year: 2026
- Tasks: Travel planning and constrained shopping
- Format: Long-horizon planning benchmark
- Difficulty: Constrained agent planning
- Paper: [DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints](https://arxiv.org/abs/2601.18137)

DeepPlanning focuses on global constrained optimization rather than local next-step reasoning. It is useful because many agents can execute short actions but still fail when they must gather information and plan coherently over a long horizon under hard constraints.

DeepPlanning is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (7 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 62.3% |
| 2 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 41.5% |
| 3 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 37.6% |
| 4 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 26.4% |
| 5 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 25.9% |
| 6 | [GLM-5](/models/glm-5) | Z.AI | 14.6% |
| 7 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 14.4% |

## FAQ

### What does DeepPlanning measure?

A long-horizon planning benchmark that tests whether agents can optimize under explicit time, budget, and feasibility constraints.

### Which model scores highest on DeepPlanning?

Qwen3.7 Plus by Alibaba currently leads with a score of 62.3% on DeepPlanning.

### How many models are evaluated on DeepPlanning?

7 AI models have been evaluated on DeepPlanning on BenchLM.

### Does DeepPlanning affect BenchLM's overall score?

Not directly. DeepPlanning is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on DeepPlanning

- [Qwen3.7 Plus vs Qwen3.6 Plus](/compare/qwen3-6-plus-vs-qwen3-7-plus)
- [Qwen3.6 Plus vs Qwen3.5 397B](/compare/qwen3-5-397b-vs-qwen3-6-plus)
- [Qwen3.5 397B vs Claude Opus 4.5](/compare/claude-opus-4-5-vs-qwen3-5-397b)
- [Claude Opus 4.5 vs Qwen3.6-35B-A3B](/compare/claude-opus-4-5-vs-qwen3-6-35b-a3b)
