# Vals Agent Poker Bench (Agent Poker Bench)

> Which model can make the most money playing poker?

Canonical page: https://benchlm.ai/benchmarks/pokeragent

- Category: [Agentic](/agentic)
- Last updated: December 23, 2025

## About Agent Poker Bench

- Year: 2026
- Tasks: Poker-playing agent trials
- Format: Accuracy score
- Difficulty: Strategic game-agent decision making
- Paper: [Agent Poker Bench](https://www.vals.ai/benchmarks/poker_agent)

BenchLM mirrors the public Vals AI Agent Poker Bench leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

Agent Poker Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (17 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 1131.83% |
| 2 | [GPT-5](https://www.vals.ai/models/openai_gpt-5-2025-08-07) | OpenAI | 1103.18% |
| 3 | [Gemini 3 Flash Preview](https://www.vals.ai/models/google_gemini-3-flash-preview) | Google | 1100.21% |
| 4 | [DeepSeek V3p2 Thinking](https://www.vals.ai/models/fireworks_deepseek-v3p2-thinking) | Fireworks | 1090.30% |
| 5 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | xAI | 1079.22% |
| 6 | [Gemini 3 Pro Preview](https://www.vals.ai/models/google_gemini-3-pro-preview) | Google | 1078.90% |
| 7 | [DeepSeek V3p1](https://www.vals.ai/models/fireworks_deepseek-v3p1) | Fireworks | 1058.23% |
| 8 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 1055.50% |
| 9 | [GPT-5.1](/models/gpt-5-1) | OpenAI | 1038.59% |
| 10 | [Grok 4 Fast (Reasoning)](/models/grok-4-fast-reasoning) | xAI | 1034.30% |
| 11 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 1033.38% |
| 12 | [Gemini 2.5 Pro](/models/gemini-2-5-pro) | Google | 1032.60% |
| 13 | [GPT-OSS 120B](/models/gpt-oss-120b) | OpenAI | 1015.33% |
| 14 | [Kimi K2 Thinking](https://www.vals.ai/models/kimi_kimi-k2-thinking) | Moonshot AI | 1011.63% |
| 15 | [Qwen3 Max Preview](https://www.vals.ai/models/alibaba_qwen3-max-preview) | Alibaba | 994.51% |
| 16 | [GLM-4.6](/models/glm-4-6) | Z.AI | 945.76% |
| 17 | [Llama4 Maverick Instruct Basic](https://www.vals.ai/models/fireworks_llama4-maverick-instruct-basic) | Fireworks | 890.50% |

## FAQ

### What does Agent Poker Bench measure?

Which model can make the most money playing poker?

### Which model leads the published Agent Poker Bench snapshot?

GPT-5.2 currently leads the published Agent Poker Bench snapshot with a score of 1131.83%.

### How many models are evaluated on Agent Poker Bench?

The December 23, 2025 contains 17 AI models.

### Does Agent Poker Bench affect BenchLM's overall score?

Not directly. Agent Poker Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
