# OpenAI MRCR v2 8-needle 128K-256K (MRCR v2 128K-256K)

> MRCR v2 slice focused on very long contexts at 128K-256K lengths.

Canonical page: https://benchlm.ai/benchmarks/mrcr-v2-128k-256k

- Category: [Reasoning](/reasoning)
- Last updated: September 18, 2026

## About MRCR v2 128K-256K

- Year: 2026
- Tasks: 8-needle retrieval tasks
- Format: Very-long-context retrieval
- Difficulty: Very long-context reasoning
- Paper: [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/)

A harder MRCR setting that stresses memory discipline and retrieval deeper into long contexts.

MRCR v2 128K-256K is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 87.5% |
| 2 | [Claude Opus 4.7 (Adaptive)](/models/claude-opus-4-7-adaptive) | Anthropic | 59.2% |

## FAQ

### What does MRCR v2 128K-256K measure?

MRCR v2 slice focused on very long contexts at 128K-256K lengths.

### Which model scores highest on MRCR v2 128K-256K?

GPT-5.5 by OpenAI currently leads with a score of 87.5% on MRCR v2 128K-256K.

### How many models are evaluated on MRCR v2 128K-256K?

2 AI models have been evaluated on MRCR v2 128K-256K on BenchLM.

### Does MRCR v2 128K-256K affect BenchLM's overall score?

Not directly. MRCR v2 128K-256K is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MRCR v2 128K-256K

- [GPT-5.5 vs Claude Opus 4.7 (Adaptive)](/compare/claude-opus-4-7-adaptive-vs-gpt-5-5)
