EdgeBench
We show this table for reference; we do not rank on it.
A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.
Score @12h on EdgeBench — July 2, 2026 release
We mirror the published score @12h view for EdgeBench. Claude Opus 4.8 leads the public snapshot at 51.3%, followed by GPT-5.5 (48.4%) and GPT-5.4 (39.3%). We do not use these results to rank models overall.
Claude Opus 4.8
Anthropic
GPT-5.5
OpenAI
GPT-5.4
OpenAI
5 modelsCodingCurrentDisplay onlyUpdated July 2, 2026 release
Score @12h table (5 models)
ScoreHow EdgeBench is shown here
BenchLM mirrors the official EdgeBench results published by ByteDance Seed with the July 2, 2026 release. EdgeBench contains 134 real-world, day-scale tasks across 6 domains, built by domain experts averaging 57.2 hours per task, and 51 tasks are publicly released together with the SForge evaluation harness.
This page ranks models by the primary published metric: average score after 12 hours of agent interaction on the full 134-task suite. The mirrored snapshot also preserves each model's score on the 51-task open-source subset.
EdgeBench is display only on BenchLM. The published rows measure long-horizon agent runs inside the SForge harness rather than normalized model-only comparisons, and most of the task suite is not public, so BenchLM does not use these scores as weighted ranking inputs.
Snapshot
The published EdgeBench snapshot places Claude Opus 4.8 first at 51.3%. The third row is 12.0 points behind. The broader top-10 range is 20.3 points, so the table still separates the published systems.
5 models have been evaluated on EdgeBench. The benchmark falls in the Coding category. EdgeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About EdgeBench
Year
2026
Tasks
Systems and software-engineering tasks
Format
Time-budgeted agent learning curves
Difficulty
Long-horizon engineering
BenchLM tracks EdgeBench as source metadata for now. The reviewed site, paper, GitHub repository, and Hugging Face dataset describe tasks and learning-curve methodology, but do not provide a stable aggregate model leaderboard suitable for scored model rows.
Freshness and provenance
Version
EdgeBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
51 of 134 tasks public with the SForge harness
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does EdgeBench measure?
A systems and software-engineering benchmark from ByteDance Seed that evaluates agents on long-horizon edge tasks using time-budgeted learning curves rather than a single static pass rate.
Which model leads the published EdgeBench snapshot?
Claude Opus 4.8 currently leads the published EdgeBench snapshot with 51.3% score @12h. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on EdgeBench?
The July 2, 2026 release snapshot contains 5 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.