Ramp Accounting Bench
We show this table for reference; we do not rank on it.
A Ramp Labs public leaderboard for long-horizon accounting agent work. BenchLM shows it as display-only reference data and excludes it from model rankings.
Mean criteria score @3 on Ramp Accounting Bench — September 29, 2026 snapshot
We mirror the published mean criteria score @3 view for Ramp Accounting Bench. Claude Opus 5.5 leads the public snapshot at 50.5%, followed by Claude Fable 5.1 (49.7%) and GPT-6 Astra (48.7%). We do not use these results to rank models overall.
Claude Opus 5.5
Anthropic
Mercor Archipelago ReAct loop · High reasoning
Pass@3 19.7% · 3 attempts per task
Claude Fable 5.1
Anthropic
Mercor Archipelago ReAct loop · High reasoning
Pass@3 19.7% · 3 attempts per task
GPT-6 Astra
OpenAI
Mercor Archipelago ReAct loop · High reasoning
Pass@3 19.0% · 3 attempts per task
24 modelsAgenticCurrentDisplay onlyUpdated September 29, 2026 snapshot
Mean criteria score @3 table (24 models)
ScoreHow to read this leaderboard
The headline is the mean criterion-level score across three independent attempts. Compare configurations as complete agent systems, not as model-only measurements.
Operator receipt: 24 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5.5 at 50.5%.
Honest limit: Ramp runs a fixed agent loop, tool environment, high reasoning effort, budget, and judge configuration. The source table is display only and does not provide model-row provenance or weighted ranking input.
How BenchLM shows Ramp Accounting Bench
BenchLM mirrors Ramp Labs’ public Ramp Accounting Bench table, captured on September 29, 2026 snapshot. The source evaluates 24 model configurations across 137 accounting tasks, with 3 independent attempts per task.
The visible score is Ramp’s mean criterion-level score across those attempts. The source’s Mercor Archipelago setup, high reasoning effort, tool environment, budget, and judge all contribute to a result, so the table is display only and excluded from BenchLM’s overall and category rankings.
Snapshot
The published Ramp Accounting Bench snapshot places Claude Opus 5.5 first at 50.5%. The third row is 1.8 points behind. The broader top-10 range is 8.1 points, so many of the published results sit in a relatively narrow band.
24 models have been evaluated on Ramp Accounting Bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Ramp Accounting Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Ramp Accounting Bench
Year
2026
Tasks
137 accounting tasks · 3 attempts per task
Format
Mean criterion-level score across three attempts
Difficulty
Long-horizon accounting agent evaluation
The published snapshot evaluates 24 model configurations on 137 accounting tasks, with 3 attempts per task.
Freshness and provenance
Version
Ramp Accounting Bench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Ramp Accounting Bench measure?
A Ramp Labs public leaderboard for long-horizon accounting agent work. BenchLM shows it as display-only reference data and excludes it from model rankings.
Which model leads the published Ramp Accounting Bench snapshot?
Claude Opus 5.5 currently leads the published Ramp Accounting Bench snapshot with 50.5% mean criteria score @3. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Ramp Accounting Bench?
The September 29, 2026 snapshot snapshot contains 24 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.