# AI-Needle

> A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.

Canonical page: https://benchlm.ai/benchmarks/aineedle

- Category: [Reasoning](/reasoning)
- Last updated: September 18, 2026

## About AI-Needle

- Year: 2026
- Tasks: Long-context retrieval
- Format: Needle-in-a-haystack recall
- Difficulty: Long-context memory
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

AI-Needle is useful for testing whether very large context windows are actually usable rather than just headline numbers. It rewards precise recall under distractors and long-document clutter.

AI-Needle is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (4 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 74% |
| 2 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 68.7% |
| 3 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 68.3% |
| 4 | [GLM-5](/models/glm-5) | Z.AI | 63.3% |

## FAQ

### What does AI-Needle measure?

A long-context retrieval benchmark that measures whether a model can recover relevant information embedded deep inside very long contexts.

### Which model scores highest on AI-Needle?

Claude Opus 4.5 by Anthropic currently leads with a score of 74% on AI-Needle.

### How many models are evaluated on AI-Needle?

4 AI models have been evaluated on AI-Needle on BenchLM.

### Does AI-Needle affect BenchLM's overall score?

Not directly. AI-Needle is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on AI-Needle

- [Claude Opus 4.5 vs Qwen3.5 397B](/compare/claude-opus-4-5-vs-qwen3-5-397b)
- [Qwen3.5 397B vs Qwen3.6 Plus](/compare/qwen3-5-397b-vs-qwen3-6-plus)
- [Qwen3.6 Plus vs GLM-5](/compare/glm-5-vs-qwen3-6-plus)
