Benchmark profile
Vals Time Horizon Index: Kerbal Space Program (Time Horizon Index: KSP)
A Vals AI agent benchmark that gives each system five days to build and run a space program in Kerbal Space Program.
Data verifiedHow BenchLM shows Time Horizon Index: KSP
BenchLM mirrors the public Vals AI Time Horizon Index: KSP leaderboard captured from https://www.vals.ai/benchmarks/time_horizon_index and updated by Vals on July 18, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Time Horizon Index: KSP is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
KSP progress on Time Horizon Index: KSP — July 18, 2026
BenchLM mirrors the published ksp progress view for Time Horizon Index: KSP. Claude Opus 4.8 leads the public snapshot at 11.833% , followed by GPT-5.5 (8.333%) and Grok 4.5 (6.333%). BenchLM does not use these results to rank models overall.
Claude Opus 4.8
Anthropic
anthropic/claude-opus-4-8
GPT-5.5
OpenAI
openai/gpt-5.5
Grok 4.5
xAI
grok/grok-4.5
KSP progress table (4 models)
ScoreThe published Time Horizon Index: KSP snapshot places Claude Opus 4.8 first at 11.833%. The third row is 5.500 points behind. The broader top-10 range is 10.166 points, so the table still separates the published systems.
4 models have been evaluated on Time Horizon Index: KSP. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Time Horizon Index: KSP is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Time Horizon Index: KSP
Year
2026
Tasks
30 progressively harder Kerbal Space Program missions
Format
Mission-ladder progress with partial credit
Difficulty
Long-horizon autonomous computer use and planning
The first Time Horizon Index uses a 30-rung mission ladder and awards partial progress within a rung. We mirror the four launch rows and progress-efficiency scores. Each result comes from one five-day run with a computer-use harness, so the small system-level table remains display only.
BenchLM freshness & provenance
Version
Time Horizon Index: KSP 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Time Horizon Index: KSP measure?
A Vals AI agent benchmark that gives each system five days to build and run a space program in Kerbal Space Program.
Which model leads the published Time Horizon Index: KSP snapshot?
Claude Opus 4.8 currently leads the published Time Horizon Index: KSP snapshot with 11.833% ksp progress. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Time Horizon Index: KSP?
4 AI models are included in BenchLM's mirrored Time Horizon Index: KSP snapshot, based on the public leaderboard captured on July 18, 2026.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.