# EdgeBench

> A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.

Canonical page: https://benchlm.ai/benchmarks/edgebench

- Category: [Coding](/coding)
- Last updated: July 2, 2026 release

## About EdgeBench

- Year: 2026
- Tasks: Systems and software-engineering tasks
- Format: Time-budgeted agent learning curves
- Difficulty: Long-horizon engineering
- Paper: [EdgeBench](https://edge-bench.org/)

BenchLM tracks EdgeBench as source metadata for now. The reviewed site, paper, GitHub repository, and Hugging Face dataset describe tasks and learning-curve methodology, but do not provide a stable aggregate model leaderboard suitable for scored model rows.

EdgeBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (5 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 51.3% |
| 2 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 48.4% |
| 3 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 39.3% |
| 4 | [GLM-5.1](/models/glm-5-1) | Z.AI | 37.4% |
| 5 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 31% |

## FAQ

### What does EdgeBench measure?

A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.

### Which model leads the published EdgeBench snapshot?

Claude Opus 4.8 currently leads the published EdgeBench snapshot with a score of 51.3%.

### How many models are evaluated on EdgeBench?

The July 2, 2026 release contains 5 AI models.

### Does EdgeBench affect BenchLM's overall score?

Not directly. EdgeBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
