Testing the Limits of Chain-of-thought with Multistep Soft Reasoning (MuSR)
A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.
Data verified 27 confirmed releases in the last 30 daysSee the free Radar BriefAbout MuSR
Year
2023
Tasks
Multi-step reasoning
Format
Narrative-based reasoning
Difficulty
Complex reasoning tasks
MuSR challenges models to perform multistep reasoning over complex narratives. Unlike simple factual questions, it requires models to track multiple entities, relationships, and logical steps across extended contexts.
BenchLM freshness & provenance
Version
MuSR 2023
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MuSR measure?
A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.
Which model scores highest on MuSR?
No models have been evaluated on MuSR yet.
How many models are evaluated on MuSR?
0 AI models have been evaluated on MuSR on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.