Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief

Testing the Limits of Chain-of-thought with Multistep Soft Reasoning (MuSR)

A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.

Data verified 27 confirmed releases in the last 30 daysSee the free Radar Brief

About MuSR

Year

2023

Tasks

Multi-step reasoning

Format

Narrative-based reasoning

Difficulty

Complex reasoning tasks

MuSR challenges models to perform multistep reasoning over complex narratives. Unlike simple factual questions, it requires models to track multiple entities, relationships, and logical steps across extended contexts.

BenchLM freshness & provenance

Version

MuSR 2023

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MuSR measure?

A dataset for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Tests the ability to perform complex, structured reasoning.

Which model scores highest on MuSR?

No models have been evaluated on MuSR yet.

How many models are evaluated on MuSR?

0 AI models have been evaluated on MuSR on BenchLM.

Last updated: September 2, 2026 · BenchLM version MuSR 2023

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.