Vals Tax Agent Bench (Tax Agent Bench)
We show this table for reference; we do not rank on it.
Agent evaluation on research-grade US tax questions using the Tax Agent Bench harness.
Overall score on Tax Agent Bench — September 22, 2026
We mirror the published overall score view for Tax Agent Bench. Claude Fable 5.1 leads the public snapshot at 77.64%, followed by Claude Opus 5 (75.06%) and GLM-5.3 (73.09%). We do not use these results to rank models overall.
Claude Fable 5.1
Anthropic
Claude Opus 5
Anthropic
GLM-5.3
Z.AI
max reasoning
26 modelsAgenticCurrentDisplay onlyUpdated September 22, 2026
Overall score table (26 models)
ScoreHow Tax Agent Bench is shown here
BenchLM mirrors the public Vals AI Tax Agent Bench leaderboard captured from https://www.vals.ai/benchmarks/tax_agent_bench and updated by Vals on September 22, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
Tax Agent Bench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Snapshot
The published Tax Agent Bench snapshot places Claude Fable 5.1 first at 77.64%. The third row is 4.55 points behind. The broader top-10 range is 11.68 points, so the table still separates the published systems.
26 models have been evaluated on Tax Agent Bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Tax Agent Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Tax Agent Bench
Year
2026
Tasks
US tax research, calculations, forms, filings, and precedent analysis
Format
Overall score; all-pass score shown separately
Difficulty
Research-grade US tax questions
The public table preserves overall and all-pass scores across fact-pattern analysis, source lookup, temporal analysis, calculations, filings, and precedent. These private-dataset agent results remain display-only and do not enter weighted model rankings.
Freshness and provenance
Version
Tax Agent Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Tax Agent Bench measure?
Agent evaluation on research-grade US tax questions using the Tax Agent Bench harness.
Which model leads the published Tax Agent Bench snapshot?
Claude Fable 5.1 currently leads the published Tax Agent Bench snapshot with 77.64% overall score. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Tax Agent Bench?
The September 22, 2026 snapshot contains 26 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.