# Ramp Accounting Bench

> A Ramp Labs public leaderboard for long-horizon accounting agent work. BenchLM shows it as display-only reference data and excludes it from model rankings.

Canonical page: https://benchlm.ai/benchmarks/ramp-accounting-bench

- Category: [Agentic](/agentic)
- Last updated: September 29, 2026 snapshot

## About Ramp Accounting Bench

- Year: 2026
- Tasks: 137 accounting tasks · 3 attempts per task
- Format: Mean criterion-level score across three attempts
- Difficulty: Long-horizon accounting agent evaluation
- Paper: [Ramp Accounting Bench](https://labs.ramp.com/ramp-accounting-bench#explore)

The published snapshot evaluates 24 model configurations on 137 accounting tasks, with 3 attempts per task.

Ramp Accounting Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (24 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 19.7% · 3 attempts per task | Anthropic | 50.53% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 19.7% · 3 attempts per task | Anthropic | 49.68% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 19.0% · 3 attempts per task | OpenAI | 48.68% |
| 4 | [Claude Opus 5](/models/claude-opus-5) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 17.5% · 3 attempts per task | Anthropic | 46.60% |
| 5 | [GPT-6.1 Sol](/models/gpt-6-1-sol) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 16.8% · 3 attempts per task | OpenAI | 46.35% |
| 6 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 14.6% · 3 attempts per task | Google | 45.74% |
| 7 | [Grok 4.7](/models/grok-4-7) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 19.0% · 3 attempts per task | xAI | 44.10% |
| 8 | [Grok 4.6](/models/grok-4-6) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 15.3% · 3 attempts per task | xAI | 43.50% |
| 9 | [GLM-5.3](/models/glm-5-3) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 21.2% · 3 attempts per task | Z.AI | 43.09% |
| 10 | [Claude Sonnet 5.5](/models/claude-sonnet-5-5) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 12.4% · 3 attempts per task | Anthropic | 42.46% |
| 11 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 15.3% · 3 attempts per task | Google | 42.30% |
| 12 | [Claude Fable 5](/models/claude-fable) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 16.1% · 3 attempts per task | Anthropic | 41.94% |
| 13 | [Qwen3.8 Max](/models/qwen3-8-max) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 19.7% · 3 attempts per task | Alibaba | 41.18% |
| 14 | [GPT-6 Sol](/models/gpt-6-sol) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 16.1% · 3 attempts per task | OpenAI | 40.91% |
| 15 | [DeepSeek V4.1 Flash](/models/deepseek-v4-1-flash) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 16.8% · 3 attempts per task | DeepSeek | 40.79% |
| 16 | [Kimi K3](/models/kimi-k3) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 13.9% · 3 attempts per task | Moonshot AI | 39.18% |
| 17 | [GLM-5.2](/models/glm-5-2) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 14.6% · 3 attempts per task | Z.AI | 37.55% |
| 18 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 12.4% · 3 attempts per task | OpenAI | 37.14% |
| 19 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 15.3% · 3 attempts per task | DeepSeek | 36.29% |
| 20 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 13.9% · 3 attempts per task | Z.AI | 35.69% |
| 21 | [GPT-6 Luna](/models/gpt-6-luna) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 10.9% · 3 attempts per task | OpenAI | 33.96% |
| 22 | [Claude Sonnet 5](/models/claude-sonnet-5) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 13.9% · 3 attempts per task | Anthropic | 33.84% |
| 23 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 8.8% · 3 attempts per task | OpenAI | 31.41% |
| 24 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | Mercor Archipelago ReAct loop · High reasoning · Pass@3 7.3% · 3 attempts per task | OpenAI | 29.01% |

## FAQ

### What does Ramp Accounting Bench measure?

A Ramp Labs public leaderboard for long-horizon accounting agent work. BenchLM shows it as display-only reference data and excludes it from model rankings.

### Which model leads the published Ramp Accounting Bench snapshot?

Claude Opus 5.5 currently leads the published Ramp Accounting Bench snapshot with a score of 50.53%.

### How many models are evaluated on Ramp Accounting Bench?

The September 29, 2026 snapshot contains 24 AI models.

### Does Ramp Accounting Bench affect BenchLM's overall score?

Not directly. Ramp Accounting Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
