# Senior SWE-Bench

> A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.

Canonical page: https://benchlm.ai/benchmarks/seniorswebench

- Category: [Coding](/coding)
- Last updated: v2026.06 release

## About Senior SWE-Bench

- Year: 2026
- Tasks: Senior-level repository tasks
- Format: Agentic software-engineering evaluation
- Difficulty: Professional senior engineering
- Paper: [Senior SWE-Bench](https://senior-swe-bench.snorkel.ai/)

BenchLM tracks Senior SWE-Bench as source metadata for now. The public page describes 50 public and 50 private tasks and renders a model comparison chart, but the crawler output does not expose exact aggregate model scores for each row.

Senior SWE-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (9 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 24% |
| 2 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 19.4% |
| 3 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 16% |
| 4 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 14.1% |
| 5 | [GLM-5.2](/models/glm-5-2) | Z.AI | 12.5% |
| 6 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 8.2% |
| 7 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 8.2% |
| 8 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 6.1% |
| 9 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 3% |

## FAQ

### What does Senior SWE-Bench measure?

A Snorkel AI benchmark of senior-level software engineering tasks emphasizing under-specified feature work, bug/performance investigation, and taste-based correctness.

### Which model leads the published Senior SWE-Bench snapshot?

Claude Opus 4.8 currently leads the published Senior SWE-Bench snapshot with a score of 24%.

### How many models are evaluated on Senior SWE-Bench?

The v2026.06 release contains 9 AI models.

### Does Senior SWE-Bench affect BenchLM's overall score?

Not directly. Senior SWE-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
