# Vals Public Benefits Bench v1.1 (Public Benefits Bench v1.1)

> Can AI help people navigate SNAP benefits?

Canonical page: https://benchlm.ai/benchmarks/publicbenefitsbench

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About Public Benefits Bench v1.1

- Year: 2026
- Tasks: SNAP public-benefits navigation tasks
- Format: Accuracy score
- Difficulty: Public-benefits policy navigation
- Paper: [Public Benefits Bench v1.1](https://www.vals.ai/benchmarks/public-benefits-bench)

BenchLM mirrors the public Vals AI Public Benefits Bench v1.1 leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

Public Benefits Bench v1.1 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (43 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 76.93% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | Anthropic | 74.90% |
| 3 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 70.64% |
| 4 | [Claude Fable 5](/models/claude-fable) | Anthropic | 70.43% |
| 5 | [MiMo-V2.6-Pro](/models/mimo-v2-6-pro) | Xiaomi | 68.94% |
| 6 | [Hy4 preview](/models/hy4-preview) | Tencent | 68.61% |
| 7 | [GLM-5.3](/models/glm-5-3) | Z.AI | 68.54% |
| 8 | [Muse Spark 1.2](/models/muse-spark-1-2) | Meta | 68.47% |
| 9 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 68.27% |
| 10 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 68.13% |
| 11 | [MiMo-V2.6-Flash](/models/mimo-v2-6-flash) | Xiaomi | 67.59% |
| 12 | [Claude Sonnet 5.5](/models/claude-sonnet-5-5) | Anthropic | 67.19% |
| 13 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 67.12% |
| 14 | [Grok 4.6](/models/grok-4-6) | xAI | 66.85% |
| 15 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 66.51% |
| 16 | [Claude Sonnet 5](/models/claude-sonnet-5) | Anthropic | 66.03% |
| 17 | [Grok 4.7](/models/grok-4-7) | xAI | 65.63% |
| 18 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Google | 65.29% |
| 19 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | DeepSeek | 64.28% |
| 20 | [MiniMax M3](/models/minimax-m3) | MiniMax | 64.14% |
| 21 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 62.92% |
| 22 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 62.45% |
| 23 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 62.38% |
| 24 | [GLM-5.1](/models/glm-5-1) | Z.AI | 61.84% |
| 25 | [Muse Spark 1.1](/models/muse-spark-1-1) | Meta | 61.30% |
| 26 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 61.16% |
| 27 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 60.89% |
| 28 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 59.47% |
| 29 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 59.41% |
| 30 | [Inkling](/models/inkling) | Thinking Machines Lab | 58.46% |
| 31 | [GPT-6 Luna](/models/gpt-6-luna) | OpenAI | 57.65% |
| 32 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 56.63% |
| 33 | [GPT-6 Sol](/models/gpt-6-sol) | OpenAI | 56.63% |
| 34 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Google | 56.56% |
| 35 | [Claude Haiku 4.5](/models/claude-haiku-4-5) | Anthropic | 54.26% |
| 36 | [Gemini 3.1 Pro Preview](https://www.vals.ai/models/google_gemini-3.1-pro-preview) | Google | 53.79% |
| 37 | [Gemini 3.5 Flash-Lite](/models/gemini-3-5-flash-lite) | Google | 51.83% |
| 38 | [Grok 4.3](/models/grok-4-3) | xAI | 51.69% |
| 39 | [Ling 3.0 Flash 2607](https://www.vals.ai/models/ant_ling-3.0-flash-2607) | ant | 50.74% |
| 40 | [Mercury 2.5](/models/mercury-2-5) | Inception | 45.26% |
| 41 | [Grok 4.1 Fast (Reasoning)](/models/grok-4-1-fast-reasoning) | xAI | 44.79% |
| 42 | [Laguna M.1](/models/laguna-m-1) | Poolside | 43.84% |
| 43 | [Laguna XS.2](/models/laguna-xs-2) | Poolside | 41.41% |

## FAQ

### What does Public Benefits Bench v1.1 measure?

Can AI help people navigate SNAP benefits?

### Which model leads the published Public Benefits Bench v1.1 snapshot?

Claude Opus 5 currently leads the published Public Benefits Bench v1.1 snapshot with a score of 76.93%.

### How many models are evaluated on Public Benefits Bench v1.1?

The September 27, 2026 contains 43 AI models.

### Does Public Benefits Bench v1.1 affect BenchLM's overall score?

Not directly. Public Benefits Bench v1.1 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
