Skip to main content
BenchLM

Vals Time Horizon Index: Kerbal Space Program (Time Horizon Index: KSP)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A Vals AI agent benchmark that gives each system five days to build and run a space program in Kerbal Space Program.

KSP progress on Time Horizon Index: KSP — September 14, 2026

We mirror the published ksp progress view for Time Horizon Index: KSP. GPT-6 Astra leads the public snapshot at 90.500%, followed by Claude Fable 5.1 (63.333%) and GPT-5.6 Sol (23.833%). We do not use these results to rank models overall.

11 modelsAgenticCurrentDisplay onlyUpdated September 14, 2026

KSP progress table (11 models)

Score
1
GPT-6 AstraOpenAI · Closedmax reasoning
90.500%
2
Claude Fable 5.1Anthropic · Closed
63.333%
3
GPT-5.6 SolOpenAI · Closedmax reasoning
23.833%
4
Gemini 3.8 FlashGoogle · Closedhigh reasoning
18.833%
5
Claude Opus 5Anthropic · Closed
18.833%
6
Claude Opus 4.8Anthropic · Closed
11.833%
7
Kimi K3Moonshot AI · Closedmax reasoning
10.500%
8
GPT-5.5OpenAI · Closedxhigh reasoning
8.333%
9
Grok 4.5xAI · Closedhigh reasoning
6.333%
10
Grok 4.6xAI · Closedhigh reasoning
5.833%
11
Muse Spark 1.1Meta · Closedxhigh reasoning
1.667%

How Time Horizon Index: KSP is shown here

BenchLM mirrors the public Vals AI Time Horizon Index: KSP leaderboard captured from https://www.vals.ai/benchmarks/time_horizon_index and updated by Vals on September 14, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

Time Horizon Index: KSP is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

11 Vals rows2 task viewsprivate datasetTasks: Overall, Progress efficiencyDisplay only

The published Time Horizon Index: KSP snapshot places GPT-6 Astra first at 90.500%. The third row is 66.667 points behind. The broader top-10 range is 84.667 points, so the table still separates the published systems.

11 models have been evaluated on Time Horizon Index: KSP. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Time Horizon Index: KSP is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Time Horizon Index: KSP

Year

2026

Tasks

30 progressively harder Kerbal Space Program missions

Format

Mission-ladder progress with partial credit

Difficulty

Long-horizon autonomous computer use and planning

The first Time Horizon Index uses a 30-rung mission ladder and awards partial progress within a rung. We mirror the four launch rows and progress-efficiency scores. Each result comes from one five-day run with a computer-use harness, so the small system-level table remains display only.

Freshness and provenance

Version

Time Horizon Index: KSP 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Time Horizon Index: KSP measure?

A Vals AI agent benchmark that gives each system five days to build and run a space program in Kerbal Space Program.

Which model leads the published Time Horizon Index: KSP snapshot?

GPT-6 Astra currently leads the published Time Horizon Index: KSP snapshot with 90.500% ksp progress. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Time Horizon Index: KSP?

The September 14, 2026 snapshot contains 11 AI models.

Last updated: September 14, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.