# Vals Time Horizon Index: Kerbal Space Program (Time Horizon Index: KSP)

> A Vals AI agent benchmark that gives each system five days to build and run a space program in Kerbal Space Program.

Canonical page: https://benchlm.ai/benchmarks/valstimehorizonksp

- Category: [Agentic](/agentic)
- Last updated: September 14, 2026

## About Time Horizon Index: KSP

- Year: 2026
- Tasks: 30 progressively harder Kerbal Space Program missions
- Format: Mission-ladder progress with partial credit
- Difficulty: Long-horizon autonomous computer use and planning
- Paper: [Time Horizon Index: KSP](https://www.vals.ai/benchmarks/time_horizon_index)

The first Time Horizon Index uses a 30-rung mission ladder and awards partial progress within a rung. We mirror the four launch rows and progress-efficiency scores. Each result comes from one five-day run with a computer-use harness, so the small system-level table remains display only.

Time Horizon Index: KSP is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (11 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | max reasoning | OpenAI | 90.500% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | — | Anthropic | 63.333% |
| 3 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | max reasoning | OpenAI | 23.833% |
| 4 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | high reasoning | Google | 18.833% |
| 5 | [Claude Opus 5](/models/claude-opus-5) | — | Anthropic | 18.833% |
| 6 | [Claude Opus 4.8](/models/claude-opus-4-8) | — | Anthropic | 11.833% |
| 7 | [Kimi K3](/models/kimi-k3) | max reasoning | Moonshot AI | 10.500% |
| 8 | [GPT-5.5](/models/gpt-5-5) | xhigh reasoning | OpenAI | 8.333% |
| 9 | [Grok 4.5](/models/grok-4-5) | high reasoning | xAI | 6.333% |
| 10 | [Grok 4.6](/models/grok-4-6) | high reasoning | xAI | 5.833% |
| 11 | [Muse Spark 1.1](/models/muse-spark-1-1) | xhigh reasoning | Meta | 1.667% |

## FAQ

### What does Time Horizon Index: KSP measure?

A Vals AI agent benchmark that gives each system five days to build and run a space program in Kerbal Space Program.

### Which model leads the published Time Horizon Index: KSP snapshot?

GPT-6 Astra currently leads the published Time Horizon Index: KSP snapshot with a score of 90.500%.

### How many models are evaluated on Time Horizon Index: KSP?

The September 14, 2026 contains 11 AI models.

### Does Time Horizon Index: KSP affect BenchLM's overall score?

Not directly. Time Horizon Index: KSP is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
