Skip to main content
BenchLM

Graphwalks BFS 0K-128K (Graphwalks BFS 128K)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Long-context graph traversal benchmark using breadth-first search tasks.

Benchmark score on Graphwalks BFS 128K — September 27, 2026

We compile the Graphwalks BFS 128K rows from provider self-reports. MAI-Thinking-1 leads the table at 90%. We do not use these results to rank models overall.

1 modelReasoningCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (1 model)

Score
1
MAI-Thinking-1Microsoft · Closed
90%

About Graphwalks BFS 128K

Year

2026

Tasks

Graph traversal tasks

Format

Long-context graph reasoning

Difficulty

Algorithmic long-context reasoning

Graphwalks BFS tests whether a model can preserve algorithmic state while traversing graph structures across long contexts.

Freshness and provenance

Version

Graphwalks BFS 128K 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Graphwalks BFS 128K measure?

Long-context graph traversal benchmark using breadth-first search tasks.

Which model scores highest on Graphwalks BFS 128K?

MAI-Thinking-1 by Microsoft currently leads with a score of 90% on Graphwalks BFS 128K.

How many models are evaluated on Graphwalks BFS 128K?

1 AI models have been evaluated on Graphwalks BFS 128K on BenchLM.

Last updated: September 27, 2026 · BenchLM version Graphwalks BFS 128K 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.