Skip to main content
BenchLM

Ramp Accounting Bench

We show this table for reference; we do not rank on it.

A Ramp Labs public leaderboard for long-horizon accounting agent work. BenchLM shows it as display-only reference data and excludes it from model rankings.

Mean criteria score @3 on Ramp Accounting Bench — September 29, 2026 snapshot

We mirror the published mean criteria score @3 view for Ramp Accounting Bench. Claude Opus 5.5 leads the public snapshot at 50.5%, followed by Claude Fable 5.1 (49.7%) and GPT-6 Astra (48.7%). We do not use these results to rank models overall.

24 modelsAgenticCurrentDisplay onlyUpdated September 29, 2026 snapshot

Mean criteria score @3 table (24 models)

Score
1
Claude Opus 5.5Anthropic · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 19.7% · 3 attempts per task
50.5%
2
Claude Fable 5.1Anthropic · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 19.7% · 3 attempts per task
49.7%
3
GPT-6 AstraOpenAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 19.0% · 3 attempts per task
48.7%
4
Claude Opus 5Anthropic · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 17.5% · 3 attempts per task
46.6%
5
GPT-6.1 SolOpenAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 16.8% · 3 attempts per task
46.4%
6
Gemini 3.8 FlashGoogle · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 14.6% · 3 attempts per task
45.7%
7
Grok 4.7xAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 19.0% · 3 attempts per task
44.1%
8
Grok 4.6xAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 15.3% · 3 attempts per task
43.5%
9
GLM-5.3Z.AI · Open weightMercor Archipelago ReAct loop · High reasoningPass@3 21.2% · 3 attempts per task
43.1%
10
Claude Sonnet 5.5Anthropic · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 12.4% · 3 attempts per task
42.5%
11
Gemini 3.7 FlashGoogle · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 15.3% · 3 attempts per task
42.3%
12
Claude Fable 5Anthropic · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 16.1% · 3 attempts per task
41.9%
13
Qwen3.8 MaxAlibaba · Open weightMercor Archipelago ReAct loop · High reasoningPass@3 19.7% · 3 attempts per task
41.2%
14
GPT-6 SolOpenAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 16.1% · 3 attempts per task
40.9%
15
DeepSeek V4.1 FlashDeepSeek · Open weightMercor Archipelago ReAct loop · High reasoningPass@3 16.8% · 3 attempts per task
40.8%
16
Kimi K3Moonshot AI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 13.9% · 3 attempts per task
39.2%
17
GLM-5.2Z.AI · Open weightMercor Archipelago ReAct loop · High reasoningPass@3 14.6% · 3 attempts per task
37.5%
18
GPT-5.6 SolOpenAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 12.4% · 3 attempts per task
37.1%
19
DeepSeek V4 Flash 0731DeepSeek · Open weightMercor Archipelago ReAct loop · High reasoningPass@3 15.3% · 3 attempts per task
36.3%
20
GLM-5.3-FlashZ.AI · Open weightMercor Archipelago ReAct loop · High reasoningPass@3 13.9% · 3 attempts per task
35.7%
21
GPT-6 LunaOpenAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 10.9% · 3 attempts per task
34.0%
22
Claude Sonnet 5Anthropic · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 13.9% · 3 attempts per task
33.8%
23
GPT-5.6 TerraOpenAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 8.8% · 3 attempts per task
31.4%
24
GPT-5.6 LunaOpenAI · ClosedMercor Archipelago ReAct loop · High reasoningPass@3 7.3% · 3 attempts per task
29.0%

How to read this leaderboard

The headline is the mean criterion-level score across three independent attempts. Compare configurations as complete agent systems, not as model-only measurements.

Operator receipt: 24 sourced rows are currently displayable on this page; the leading published row is Claude Opus 5.5 at 50.5%.

Honest limit: Ramp runs a fixed agent loop, tool environment, high reasoning effort, budget, and judge configuration. The source table is display only and does not provide model-row provenance or weighted ranking input.

How BenchLM shows Ramp Accounting Bench

BenchLM mirrors Ramp Labs’ public Ramp Accounting Bench table, captured on September 29, 2026 snapshot. The source evaluates 24 model configurations across 137 accounting tasks, with 3 independent attempts per task.

The visible score is Ramp’s mean criterion-level score across those attempts. The source’s Mercor Archipelago setup, high reasoning effort, tool environment, budget, and judge all contribute to a result, so the table is display only and excluded from BenchLM’s overall and category rankings.

Snapshot

24 model configurations137 tasks3 attempts per task9,864 agent rolloutsDisplay only

The published Ramp Accounting Bench snapshot places Claude Opus 5.5 first at 50.5%. The third row is 1.8 points behind. The broader top-10 range is 8.1 points, so many of the published results sit in a relatively narrow band.

24 models have been evaluated on Ramp Accounting Bench. The benchmark falls in the Agentic category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. Ramp Accounting Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About Ramp Accounting Bench

Year

2026

Tasks

137 accounting tasks · 3 attempts per task

Format

Mean criterion-level score across three attempts

Difficulty

Long-horizon accounting agent evaluation

The published snapshot evaluates 24 model configurations on 137 accounting tasks, with 3 attempts per task.

Freshness and provenance

Version

Ramp Accounting Bench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does Ramp Accounting Bench measure?

A Ramp Labs public leaderboard for long-horizon accounting agent work. BenchLM shows it as display-only reference data and excludes it from model rankings.

Which model leads the published Ramp Accounting Bench snapshot?

Claude Opus 5.5 currently leads the published Ramp Accounting Bench snapshot with 50.5% mean criteria score @3. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on Ramp Accounting Bench?

The September 29, 2026 snapshot snapshot contains 24 AI models.

Last updated: September 29, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.