# PinchBench

> An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.

Canonical page: https://benchlm.ai/benchmarks/pinchbench

- Category: [Agentic](/agentic)
- Last updated: 08/18/2026, 1:07 PM

## About PinchBench

- Year: 2026
- Tasks: 23 OpenClaw agent tasks
- Format: Average success rate from official runs
- Difficulty: Long-horizon agent workflows
- Paper: [About PinchBench](https://pinchbench.com/about)

PinchBench publishes official OpenClaw runs across 23 tasks and grades results with automated checks plus an LLM judge. BenchLM mirrors the public average-score view as a display-only benchmark.

PinchBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (59 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [anthropic/claude-opus-4.8-fast](https://pinchbench.com/?score=average) | Anthropic | 93.5% |
| 2 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 92.5% |
| 3 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 90.5% |
| 4 | [nvidia/nemotron-3-ultra-550b-a55b](https://pinchbench.com/?score=average) | nvidia | 89.9% |
| 5 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 89.7% |
| 6 | [Grok Build 0.1](/models/grok-build-0-1) | xAI | 88.9% |
| 7 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 88.7% |
| 8 | [qwen/qwen3.6-flash](https://pinchbench.com/?score=average) | Alibaba | 88.1% |
| 9 | [MiMo-V2.5-Pro](/models/mimo-v2-5-pro) | Xiaomi | 87.5% |
| 10 | [GLM-5.2](/models/glm-5-2) | Z.AI | 87.0% |
| 11 | [nvidia/nemotron-3.5-lightning-30b-a3b](https://pinchbench.com/?score=average) | nvidia | 86.4% |
| 12 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 84.2% |
| 13 | [inclusionai/ling-2.6-1t](https://pinchbench.com/?score=average) | inclusionai | 82.6% |
| 14 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 81.7% |
| 15 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 81.0% |
| 16 | [Gemini 3.1 Flash-Lite](/models/gemini-3-1-flash-lite) | Google | 80.5% |
| 17 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 80.3% |
| 18 | [Step 3.5 Flash](/models/step-3-5-flash) | StepFun | 79.4% |
| 19 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 79.2% |
| 20 | [Kimi K2.7 Code](/models/kimi-k2-7-code) | Moonshot AI | 76.1% |
| 21 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 76.0% |
| 22 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 75.9% |
| 23 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 75.7% |
| 24 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 75.5% |
| 25 | [Grok 4.5](/models/grok-4-5) | xAI | 75.2% |
| 26 | [Seed-2.0-Lite](/models/seed-2-0-lite) | ByteDance | 75.0% |
| 27 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 74.2% |
| 28 | [Grok 4.3](/models/grok-4-3) | xAI | 73.7% |
| 29 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 72.5% |
| 30 | [Gemini 3 Flash](/models/gemini-3-flash) | Google | 72.1% |
| 31 | [GLM-5-Turbo](/models/glm-5-turbo) | Z.AI | 71.8% |
| 32 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 69.9% |
| 33 | [sakana/fugu-ultra](https://pinchbench.com/?score=average) | sakana | 69.5% |
| 34 | [mistralai/devstral-2512](https://pinchbench.com/?score=average) | Mistral | 69.4% |
| 35 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 69.0% |
| 36 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 67.7% |
| 37 | [GLM-5V-Turbo](/models/glm-5v-turbo) | Z.AI | 67.6% |
| 38 | [MiniMax M2.7](/models/minimax-m2-7) | MiniMax | 66.8% |
| 39 | [Trinity-Large-Preview](/models/trinity-large-preview) | Arcee AI | 65.7% |
| 40 | [mistralai/mistral-small-2603](https://pinchbench.com/?score=average) | Mistral | 64.6% |
| 41 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 62.7% |
| 42 | [aion-labs/aion-3.0](https://pinchbench.com/?score=average) | aion-labs | 61.1% |
| 43 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 61.1% |
| 44 | [GLM-5.1](/models/glm-5-1) | Z.AI | 59.9% |
| 45 | [google/gemma-4-26b-a4b-it](https://pinchbench.com/?score=average) | Google | 56.4% |
| 46 | [Claude Fable 5](/models/claude-fable) | Anthropic | 54.8% |
| 47 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 54.6% |
| 48 | [mistralai/mistral-large-2512](https://pinchbench.com/?score=average) | Mistral | 54.5% |
| 49 | [Trinity-Large-Preview](/models/trinity-large-preview) | Arcee AI | 53.0% |
| 50 | [google/gemma-4-31b-it](https://pinchbench.com/?score=average) | Google | 52.7% |
| 51 | [anthropic/claude-sonnet-4](https://pinchbench.com/?score=average) | Anthropic | 48.8% |
| 52 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 44.8% |
| 53 | [Nemotron 3 Super 120B A12B](/models/nemotron-3-super-120b-a12b) | NVIDIA | 42.2% |
| 54 | [Mercury 2](/models/mercury-2) | Inception | 39.6% |
| 55 | [amazon/nova-2-lite-v1](https://pinchbench.com/?score=average) | amazon | 37.6% |
| 56 | [GPT-OSS 20B](/models/gpt-oss-20b) | OpenAI | 36.3% |
| 57 | [GPT-5.5 Pro](/models/gpt-5-5-pro) | OpenAI | 21.4% |
| 58 | [Llama 3.1 70B Instruct](https://pinchbench.com/?score=average) | Meta | 10.7% |
| 59 | [Llama 4 Scout](/models/llama-4-scout) | Meta | 3.2% |

## FAQ

### What does PinchBench measure?

An OpenClaw agent benchmark from Kilo that measures successful task completion across standardized real-world agent workflows.

### Which model leads the published PinchBench snapshot?

anthropic/claude-opus-4.8-fast currently leads the published PinchBench snapshot with a score of 93.5%.

### How many models are evaluated on PinchBench?

The 08/18/2026, 1:07 PM contains 59 AI models.

### Does PinchBench affect BenchLM's overall score?

Not directly. PinchBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
