# Vals Tax Agent Bench (Tax Agent Bench)

> Agent evaluation on research-grade US tax questions using the Tax Agent Bench harness.

Canonical page: https://benchlm.ai/benchmarks/valstaxagentbench

- Category: [Agentic](/agentic)
- Last updated: September 22, 2026

## About Tax Agent Bench

- Year: 2026
- Tasks: US tax research, calculations, forms, filings, and precedent analysis
- Format: Overall score; all-pass score shown separately
- Difficulty: Research-grade US tax questions
- Paper: [Tax Agent Bench](https://www.vals.ai/benchmarks/tax_agent_bench)

The public table preserves overall and all-pass scores across fact-pattern analysis, source lookup, temporal analysis, calculations, filings, and precedent. These private-dataset agent results remain display-only and do not enter weighted model rankings.

Tax Agent Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (26 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | — | Anthropic | 77.64% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | — | Anthropic | 75.06% |
| 3 | [GLM-5.3](/models/glm-5-3) | max reasoning | Z.AI | 73.09% |
| 4 | [Muse Spark 1.3](/models/muse-spark-1-3) | xhigh reasoning | Meta | 71.93% |
| 5 | [Grok 4.6](/models/grok-4-6) | high reasoning | xAI | 70.79% |
| 6 | [Claude Opus 5.5](/models/claude-opus-5-5) | — | Anthropic | 70.50% |
| 7 | [Kimi K3](/models/kimi-k3) | max reasoning | Moonshot AI | 68.67% |
| 8 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | max reasoning | OpenAI | 67.95% |
| 9 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | high reasoning | Google | 66.77% |
| 10 | [Qwen3.8 Max](/models/qwen3-8-max) | — | Alibaba | 65.96% |
| 11 | [Grok 4.7](/models/grok-4-7) | xhigh reasoning | xAI | 65.60% |
| 12 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | max reasoning | OpenAI | 65.20% |
| 13 | [Hy4 preview](/models/hy4-preview) | — | Tencent | 63.71% |
| 14 | [GPT-6 Astra](/models/gpt-6-astra) | max reasoning | OpenAI | 63.34% |
| 15 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | high reasoning | DeepSeek | 62.46% |
| 16 | [Claude Sonnet 5](/models/claude-sonnet-5) | — | Anthropic | 62.27% |
| 17 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | max reasoning | OpenAI | 60.81% |
| 18 | [GPT-5.5](/models/gpt-5-5) | xhigh reasoning | OpenAI | 60.46% |
| 19 | [GPT-6 Luna](/models/gpt-6-luna) | max reasoning | OpenAI | 58.86% |
| 20 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | max reasoning | DeepSeek | 58.66% |
| 21 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | high reasoning | Google | 57.66% |
| 22 | [GPT-6 Sol](/models/gpt-6-sol) | max reasoning | OpenAI | 53.05% |
| 23 | [MiniMax M3](/models/minimax-m3) | — | MiniMax | 49.69% |
| 24 | [Ling 3.0 Flash Af Rc3](https://www.vals.ai/models/ant_ling-3.0-flash-af-rc3) | — | Ant | 45.70% |
| 25 | [Inkling](/models/inkling) | 0.99 reasoning | Thinking Machines Lab | 43.52% |
| 26 | [Mercury 2.5](/models/mercury-2-5) | high reasoning | Inception | 12.77% |

## FAQ

### What does Tax Agent Bench measure?

Agent evaluation on research-grade US tax questions using the Tax Agent Bench harness.

### Which model leads the published Tax Agent Bench snapshot?

Claude Fable 5.1 currently leads the published Tax Agent Bench snapshot with a score of 77.64%.

### How many models are evaluated on Tax Agent Bench?

The September 22, 2026 contains 26 AI models.

### Does Tax Agent Bench affect BenchLM's overall score?

Not directly. Tax Agent Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
