Skip to main content
BenchLM

Vending-Bench 2

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Tests long-horizon coherence by asking agents to operate a simulated vending-machine business for one year.

Arithmetic mean final bank balance on Vending-Bench 2 — September 29, 2026

We mirror the published arithmetic mean final bank balance view for Vending-Bench 2. GPT-6 Astra leads the public snapshot at $15,514.70, followed by GPT-6 Sol ($14,427.85) and Claude Opus 5 ($11,181.87). We do not use these results to rank models overall.

66 modelsAgenticCurrentDisplay onlyUpdated September 29, 2026

Arithmetic mean final bank balance table (66 models)

Score
1
GPT-6 AstraOpenAI · Closed6 runs · SEM $1,074.48
$15,514.70
2
GPT-6 SolOpenAI · Closed6 runs · SEM $1,050.68
$14,427.85
3
Claude Opus 5Anthropic · Closed6 runs · SEM $2,093.58
$11,181.87
4
Claude Opus 4.7Anthropic · Closed6 runs · SEM $1,181.31
$10,936.76
5
Grok 4.7xAI · Closed6 runs · SEM $651.73
$10,536.83
6
GPT-5.6 SolOpenAI · Closed5 runs · SEM $1,337.80
$9,619.37
7
Claude Opus 5.5Anthropic · Closed6 runs · SEM $784.72
$9,235.25
8
Grok 4.6xAI · Closed4 runs · SEM $1,603.96
$9,047.03
9
GLM-5.2Z.AI · Open weight6 runs · SEM $1,084.11
$8,313.78
10
GLM-5.3Z.AI · Open weight6 runs · SEM $786.92
$8,163.61
11
Claude Opus 4.6Anthropic · Closed5 runs · SEM $1,366.99
$8,017.59
12
GPT-5.5OpenAI · Closed5 runs · SEM $1,345.98
$7,523.84
13
GPT-5.6 TerraOpenAI · Closed5 runs · SEM $373.27
$7,343.21
14
Claude Sonnet 4.6Anthropic · Closed5 runs · SEM $721.98
$7,204.14
15
Muse Spark 1.1Meta · Closed6 runs · SEM $1,016.33
$6,520.47
16
Claude Sonnet 5Anthropic · Closed6 runs · SEM $783.74
$6,377.70
17
Kimi K2.6Moonshot AI · Open weight5 runs · SEM $660.87
$6,204.57
18
GPT-5.4OpenAI · Closed5 runs · SEM $871.10
$6,144.18
19
GPT-5.3 CodexOpenAI · Closed5 runs · SEM $795.98
$5,940.12
20
Claude Opus 4.8Anthropic · ClosedHigh reasoning · 5 runs · SEM $830.67
$5,787.43
21
Claude Fable 5Anthropic · ClosedHigh reasoning · 5 runs · SEM $1,074.93
$5,680.26
22
GLM-5.1Z.AI · Open weight5 runs · SEM $333.62
$5,634.41
23
Gemini 3 ProGoogle · Closed5 runs · SEM $905.03
$5,478.16
24
Claude Fable 5.1Anthropic · Closed6 runs · SEM $1,244.44
$5,421.56
25
Gemini 3.5 FlashGoogle · Closed6 runs · SEM $1,058.73
$5,396.42
26
Kimi K3Moonshot AI · ClosedMoonshot endpoint · 6 runs · SEM $691.16
$5,165.04
27
Qwen 3.6 PlusAlibaba5 runs · SEM $1,097.12
$5,114.87
28
Gemini 3.8 FlashGoogle · Closed6 runs · SEM $490.16
$5,093.79
29
Kimi K2.7 CodeMoonshot AI · Open weight5 runs · SEM $1,069.37
$5,082.94
30
Claude Fable 5Anthropic · ClosedLow reasoning · 5 runs · SEM $1,225.23
$5,018.52
31
Claude Opus 4.5Anthropic · Closed6 runs · SEM $979.25
$4,967.06
32
Claude Fable 5Anthropic · ClosedMax reasoning · 5 runs · SEM $953.63
$4,966.64
33
Kimi K3Moonshot AI · ClosedFireworks endpoint · 6 runs · SEM $1,652.42
$4,906.96
34
Grok 4.20xAI · Closed6 runs · SEM $1,276.33
$4,662.85
35
Claude Fable 5Anthropic · ClosedNone reasoning · 5 runs · SEM $662.39
$4,529.94
36
GLM-5Z.AI · Open weight5 runs · SEM $393.55
$4,432.12
37
Claude Fable 5Anthropic · ClosedMedium reasoning · 4 runs · SEM $1,426.83
$4,339.81
38
Qwen 3.6 Max (preview)Alibaba · Closed6 runs · SEM $1,469.56
$4,254.19
39
GPT-5.6 LunaOpenAI · Closed5 runs · SEM $1,344.16
$4,094.71
40
Grok 4.5xAI · Closed5 runs · SEM $435.28
$3,887.43
41
Claude Sonnet 4.5Anthropic · Closed5 runs · SEM $1,102.10
$3,838.74
42
Gemini 3.1 ProGoogle · ClosedCustom tools · 5 runs · SEM $1,241.57
$3,774.25
43
Gemini 3 FlashGoogle · Closed5 runs · SEM $639.34
$3,634.72
44
GPT-5.2OpenAI · Closed5 runs · SEM $392.08
$3,591.33
45
DeepSeek V4 Pro 0813DeepSeek · Open weight6 runs · SEM $570.70
$3,284.52
46
Claude Opus 4.8Anthropic · ClosedMax reasoning · 5 runs · SEM $722.02
$2,992.34
47
GLM-4.7Z.AI · Open weight5 runs · SEM $620.60
$2,376.82
48
MiniMax M3MiniMax · Open weight6 runs · SEM $1,151.04
$2,157.77
49
GPT-5.1OpenAI · Closed5 runs · SEM $756.17
$1,473.43
50
Kimi K2.5Moonshot AI · Open weight5 runs · SEM $596.70
$1,198.46
51
Grok 4.1 FastxAI · Closed5 runs · SEM $457.65
$1,106.63
52
DeepSeek V3.2DeepSeek · Open weight5 runs · SEM $503.15
$1,034.00
53
Gemini 3.1 ProGoogle · Closed5 runs · SEM $652.55
$911.21
54
Gemini 2.5 ProGoogle · Closed5 runs · SEM $465.30
$573.64
55
Gemini 2.5 FlashGoogle · Closed5 runs · SEM $365.69
$548.84
56
Qwen 3.5 35B A3BAlibaba4 runs · SEM $486.47
$462.69
57
Claude Haiku 4.5Anthropic · Closed5 runs · SEM $274.94
$458.89
58
Qwen 3.5 27BAlibaba5 runs · SEM $134.55
$201.98
59
MiniMax-M2MiniMax5 runs · SEM $97.40
$160.60
60
Qwen3 MaxAlibaba · Closed6 runs · SEM $94.52
$71.57
61
Grok 4.3xAI · Closed6 runs · SEM $58.90
$35.26
62
Qwen 3.5 PlusAlibaba5 runs · SEM $20.31
$0.54
63
Qwen3 235B A22B ThinkingAlibaba5 runs · SEM $14.59
$-11.34
64
GPT-OSS 120BOpenAI · Open weight5 runs · SEM $2.87
$-21.53
65
MiniMax M2.5MiniMax · Closed5 runs · SEM $0.28
$-23.16
66
GPT-5 miniOpenAI · Closed5 runs · SEM $5.97
$-31.18

How we show Vending-Bench 2

We mirror Andon Labs' arithmetic-mean Vending-Bench 2 table. The current source contains 66 model-configuration rows and 349 completed runs, with 4 to 6 runs behind each row. The table ranks final bank balance after a one-year simulation and keeps each row's standard error attached.

Agents start with $500, source products, negotiate with suppliers, stock a simulated vending machine, and handle sales and customer messages. The snapshot also retains the publisher's geometric mean, but this page follows the source table's default arithmetic-mean ordering. Reasoning settings, custom-tool runs, and provider endpoints remain separate configurations.

Snapshot

66 model-configuration rows349 completed runs365-day simulationArithmetic meanDisplay only

The published Vending-Bench 2 snapshot places GPT-6 Astra first at $15,514.70. The third row is 4332.83 score units behind. The broader top-10 range is 7351.09 score units, so the table still separates the published systems.

66 models have been evaluated on Vending-Bench 2. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vending-Bench 2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Vending-Bench 2

Year

2026

Tasks

One year of simulated autonomous business operation per run

Format

Arithmetic mean final bank balance in USD

Difficulty

Long-horizon agent coherence and business operation

Agents begin with $500 and manage product sourcing, supplier negotiation, inventory, pricing, sales, and customer requests over a 365-day simulation. The leaderboard uses arithmetic mean final bank balance by default and publishes row-level standard errors. We keep model configurations separate and exclude the benchmark from weighted rankings.

Freshness and provenance

Version

Vending-Bench 2 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Vending-Bench 2 measure?

Tests long-horizon coherence by asking agents to operate a simulated vending-machine business for one year.

Which model leads the published Vending-Bench 2 snapshot?

GPT-6 Astra currently leads the published Vending-Bench 2 snapshot with $15,514.70 arithmetic mean final bank balance. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Vending-Bench 2?

The September 29, 2026 snapshot contains 66 AI models.

Last updated: September 29, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.