# Vals Public Benefits Bench v1 (Public Benefits Bench v1)

> Can AI help people navigate SNAP benefits?

Canonical page: https://benchlm.ai/benchmarks/publicbenefitsbenchv1

- Category: [Agentic](/agentic)
- Last updated: June 9, 2026

## About Public Benefits Bench v1

- Year: 2026
- Tasks: SNAP public-benefits navigation tasks
- Format: Accuracy score
- Difficulty: Public-benefits policy navigation
- Paper: [Public Benefits Bench v1](https://www.vals.ai/benchmarks/public-benefits-bench-v1)

BenchLM mirrors the public Vals AI Public Benefits Bench v1 leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

Public Benefits Bench v1 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (13 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Fable 5](/models/claude-fable) | — | Anthropic | 71.65% |
| 2 | [Claude Opus 4.8](/models/claude-opus-4-8) | — | Anthropic | 62.11% |
| 3 | [MiniMax M3](/models/minimax-m3) | — | MiniMax | 60.69% |
| 4 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | — | Anthropic | 58.52% |
| 5 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | high reasoning | Google | 57.98% |
| 6 | [GLM-5.1](/models/glm-5-1) | — | Z.AI | 57.92% |
| 7 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | max reasoning | DeepSeek | 57.58% |
| 8 | [GPT-5.5](/models/gpt-5-5) | xhigh reasoning | OpenAI | 57.24% |
| 9 | [Kimi K2.6](/models/kimi-2-6) | — | Moonshot AI | 53.38% |
| 10 | [Gemini 3.1 Pro Preview](https://www.vals.ai/models/google_gemini-3.1-pro-preview) | high reasoning | Google | 50.88% |
| 11 | [Grok 4.3](/models/grok-4-3) | high reasoning | xAI | 50.07% |
| 12 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | — | Anthropic | 49.53% |
| 13 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | — | xAI | 44.32% |

## FAQ

### What does Public Benefits Bench v1 measure?

Can AI help people navigate SNAP benefits?

### Which model leads the published Public Benefits Bench v1 snapshot?

Claude Fable 5 currently leads the published Public Benefits Bench v1 snapshot with a score of 71.65%.

### How many models are evaluated on Public Benefits Bench v1?

The June 9, 2026 contains 13 AI models.

### Does Public Benefits Bench v1 affect BenchLM's overall score?

Not directly. Public Benefits Bench v1 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
