Vending-Bench 2
Tests long-horizon coherence by asking agents to operate a simulated vending-machine business for one year.
How we show Vending-Bench 2
We mirror Andon Labs' arithmetic-mean Vending-Bench 2 table. The current source contains 59 model-configuration rows and 307 completed runs, with 4 to 6 runs behind each row. The table ranks final bank balance after a one-year simulation and keeps each row's standard error attached.
Agents start with $500, source products, negotiate with suppliers, stock a simulated vending machine, and handle sales and customer messages. The snapshot also retains the publisher's geometric mean, but this page follows the source table's default arithmetic-mean ordering. Reasoning settings, custom-tool runs, and provider endpoints remain separate configurations.
Arithmetic mean final bank balance on Vending-Bench 2 — August 14, 2026
BenchLM mirrors the published arithmetic mean final bank balance view for Vending-Bench 2. Claude Opus 5 leads the public snapshot at $11,181.87 , followed by Claude Opus 4.7 ($10,936.76) and GPT-5.6 Sol ($9,619.37). We do not use these results to rank models overall.
Claude Opus 5
Anthropic
6 runs · SEM $2,093.58
claude-opus-5
Claude Opus 4.7
Anthropic
6 runs · SEM $1,181.31
claude-opus-4-7
GPT-5.6 Sol
OpenAI
5 runs · SEM $1,337.80
gpt-5-6-sol
Arithmetic mean final bank balance table (59 models)
ScoreThe published Vending-Bench 2 snapshot places Claude Opus 5 first at $11,181.87. The third row is 1562.50 dollars behind. The broader top-10 range is 4661.40 dollars, so the table still separates the published systems.
59 models have been evaluated on Vending-Bench 2. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vending-Bench 2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Vending-Bench 2
Year
2026
Tasks
One year of simulated autonomous business operation per run
Format
Arithmetic mean final bank balance in USD
Difficulty
Long-horizon agent coherence and business operation
Agents begin with $500 and manage product sourcing, supplier negotiation, inventory, pricing, sales, and customer requests over a 365-day simulation. The leaderboard uses arithmetic mean final bank balance by default and publishes row-level standard errors. We keep model configurations separate and exclude the benchmark from weighted rankings.
BenchLM freshness & provenance
Version
Vending-Bench 2 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Vending-Bench 2 measure?
Tests long-horizon coherence by asking agents to operate a simulated vending-machine business for one year.
Which model leads the published Vending-Bench 2 snapshot?
Claude Opus 5 currently leads the published Vending-Bench 2 snapshot with $11,181.87 arithmetic mean final bank balance. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Vending-Bench 2?
59 AI models are included in BenchLM's mirrored Vending-Bench 2 snapshot, based on the public leaderboard captured on August 14, 2026.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.