Vending-Bench 2
We show this table for reference; we do not rank on it.
Tests long-horizon coherence by asking agents to operate a simulated vending-machine business for one year.
Arithmetic mean final bank balance on Vending-Bench 2 — September 29, 2026
We mirror the published arithmetic mean final bank balance view for Vending-Bench 2. GPT-6 Astra leads the public snapshot at $15,514.70, followed by GPT-6 Sol ($14,427.85) and Claude Opus 5 ($11,181.87). We do not use these results to rank models overall.
GPT-6 Astra
OpenAI
6 runs · SEM $1,074.48
GPT-6 Sol
OpenAI
6 runs · SEM $1,050.68
Claude Opus 5
Anthropic
6 runs · SEM $2,093.58
66 modelsAgenticCurrentDisplay onlyUpdated September 29, 2026
Arithmetic mean final bank balance table (66 models)
ScoreHow we show Vending-Bench 2
We mirror Andon Labs' arithmetic-mean Vending-Bench 2 table. The current source contains 66 model-configuration rows and 349 completed runs, with 4 to 6 runs behind each row. The table ranks final bank balance after a one-year simulation and keeps each row's standard error attached.
Agents start with $500, source products, negotiate with suppliers, stock a simulated vending machine, and handle sales and customer messages. The snapshot also retains the publisher's geometric mean, but this page follows the source table's default arithmetic-mean ordering. Reasoning settings, custom-tool runs, and provider endpoints remain separate configurations.
Snapshot
The published Vending-Bench 2 snapshot places GPT-6 Astra first at $15,514.70. The third row is 4332.83 score units behind. The broader top-10 range is 7351.09 score units, so the table still separates the published systems.
66 models have been evaluated on Vending-Bench 2. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Vending-Bench 2 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Vending-Bench 2
Year
2026
Tasks
One year of simulated autonomous business operation per run
Format
Arithmetic mean final bank balance in USD
Difficulty
Long-horizon agent coherence and business operation
Agents begin with $500 and manage product sourcing, supplier negotiation, inventory, pricing, sales, and customer requests over a 365-day simulation. The leaderboard uses arithmetic mean final bank balance by default and publishes row-level standard errors. We keep model configurations separate and exclude the benchmark from weighted rankings.
Freshness and provenance
Version
Vending-Bench 2 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Vending-Bench 2 measure?
Tests long-horizon coherence by asking agents to operate a simulated vending-machine business for one year.
Which model leads the published Vending-Bench 2 snapshot?
GPT-6 Astra currently leads the published Vending-Bench 2 snapshot with $15,514.70 arithmetic mean final bank balance. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Vending-Bench 2?
The September 29, 2026 snapshot contains 66 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.