# SWE-Marathon

> A long-horizon software engineering benchmark from Abundant AI with multi-hour tasks spanning library reproductions, full-stack product clones, and ML engineering.

Canonical page: https://benchlm.ai/benchmarks/swemarathon

- Category: [Coding](/coding)
- Last updated: SWE-Marathon v1.0

## About SWE-Marathon

- Year: 2026
- Tasks: 20 multi-hour software engineering tasks
- Format: Task resolution and trajectory review
- Difficulty: Ultra-long-horizon software engineering
- Paper: [SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?](https://www.swe-marathon.org/)

BenchLM tracks SWE-Marathon as a display-only external benchmark. The official v1.0 site reports 20 multi-hour tasks, 1,300 logged trials, task-level leaderboards, and replayable trajectory artifacts; BenchLM keeps it source-metadata-only until there is a stable public aggregate feed.

SWE-Marathon is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does SWE-Marathon measure?

A long-horizon software engineering benchmark from Abundant AI with multi-hour tasks spanning library reproductions, full-stack product clones, and ML engineering.

### Which model leads the published SWE-Marathon snapshot?

No models have been evaluated on SWE-Marathon yet.

### How many models are evaluated on SWE-Marathon?

The SWE-Marathon v1.0 contains 0 AI models.

### Does SWE-Marathon affect BenchLM's overall score?

Not directly. SWE-Marathon is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
