OpenAI MRCR v2 8-needle 256K-512K (MRCR v2 256K-512K)
MRCR v2 slice focused on retrieval across 256K-512K-token contexts.
Data verified 27 confirmed releases in the last 30 daysSee the free Radar BriefBenchmark score on MRCR v2 256K-512K — September 2, 2026
We mirror the published score view for MRCR v2 256K-512K. Muse Spark 1.3 leads the public snapshot at 98.5%. We do not use these results to rank models overall.
1 modelReasoningCurrentDisplay onlyUpdated September 2, 2026
Benchmark score table (1 model)
ScoreAbout MRCR v2 256K-512K
Year
2026
Tasks
100 eight-needle retrieval examples
Format
Mean sequence-matcher ratio
Difficulty
Very long-context retrieval
Meta re-bins OpenAI's MRCR v2 data by o200k_base token count and evaluates 100 examples with eight needles. The mean sequence-matcher ratio stays display-only because this slice is not part of the weighted reasoning schema.
BenchLM freshness & provenance
Version
MRCR v2 256K-512K 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MRCR v2 256K-512K measure?
MRCR v2 slice focused on retrieval across 256K-512K-token contexts.
Which model scores highest on MRCR v2 256K-512K?
Muse Spark 1.3 by Meta currently leads with a score of 98.5% on MRCR v2 256K-512K.
How many models are evaluated on MRCR v2 256K-512K?
1 AI models have been evaluated on MRCR v2 256K-512K on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.