# SWE-bench Verified (mini-swe-agent-v2) (SWE-bench Verified*)

> A display-only SWE-bench Verified reference from Arcee AI's Trinity-Large-Thinking comparison chart.

Canonical page: https://benchlm.ai/benchmarks/sweverifiedarcee

- Category: [Coding](/coding)
- Last updated: September 27, 2026

## About SWE-bench Verified*

- Year: 2026
- Tasks: Repository task completion
- Format: Agent scaffold benchmark
- Difficulty: Professional software engineering
- Paper: [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking)

BenchLM stores this chart-specific SWE-bench Verified row separately because Arcee notes all models were evaluated in mini-swe-agent-v2.

SWE-bench Verified* is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (5 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 75.6% |
| 2 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 75.4% |
| 3 | [GLM-5](/models/glm-5) | Z.AI | 72.8% |
| 4 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 70.8% |
| 5 | [Trinity-Large-Thinking](/models/trinity-large-thinking) | Arcee AI | 63.2% |

## FAQ

### What does SWE-bench Verified* measure?

A display-only SWE-bench Verified reference from Arcee AI's Trinity-Large-Thinking comparison chart.

### Which model scores highest on SWE-bench Verified*?

Claude Opus 4.6 by Anthropic currently leads with a score of 75.6% on SWE-bench Verified*.

### How many models are evaluated on SWE-bench Verified*?

5 AI models have been evaluated on SWE-bench Verified* on BenchLM.

### Does SWE-bench Verified* affect BenchLM's overall score?

Not directly. SWE-bench Verified* is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on SWE-bench Verified*

- [Claude Opus 4.6 vs MiniMax M2.7](/compare/claude-opus-4-6-vs-minimax-m2-7)
- [MiniMax M2.7 vs GLM-5](/compare/glm-5-vs-minimax-m2-7)
- [GLM-5 vs Kimi K2.5](/compare/glm-5-vs-kimi-k2-5)
- [Kimi K2.5 vs Trinity-Large-Thinking](/compare/kimi-k2-5-vs-trinity-large-thinking)
