# OpenAI MRCR v2 8-needle 64K-128K (MRCR v2 64K-128K)

> MRCR v2 slice focused on long-context retrieval at 64K-128K lengths.

Canonical page: https://benchlm.ai/benchmarks/mrcr-v2-64k-128k

- Category: [Reasoning](/reasoning)
- Last updated: September 27, 2026

## About MRCR v2 64K-128K

- Year: 2026
- Tasks: 8-needle retrieval tasks
- Format: Long-context retrieval
- Difficulty: Long-context reasoning
- Paper: [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/)

Measures whether models can recover the right details when multiple relevant items are buried in long contexts.

MRCR v2 64K-128K is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Google | 97% |
| 2 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 83.1% |

## FAQ

### What does MRCR v2 64K-128K measure?

MRCR v2 slice focused on long-context retrieval at 64K-128K lengths.

### Which model scores highest on MRCR v2 64K-128K?

Gemini 3.7 Flash by Google currently leads with a score of 97% on MRCR v2 64K-128K.

### How many models are evaluated on MRCR v2 64K-128K?

2 AI models have been evaluated on MRCR v2 64K-128K on BenchLM.

### Does MRCR v2 64K-128K affect BenchLM's overall score?

Not directly. MRCR v2 64K-128K is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MRCR v2 64K-128K

- [Gemini 3.7 Flash vs GPT-5.5](/compare/gemini-3-7-flash-vs-gpt-5-5)
