# ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job (ApprenticeBench)

> Tests whether a computer-use agent can learn a real accounts-payable job on the job, processing 100 vendor bills in a company ERP system with only the handbook, historical records, and mentor feedback a new hire would get.

Canonical page: https://benchlm.ai/benchmarks/apprenticebench

- Category: [Agentic](/agentic)
- Last updated: September 18, 2026 snapshot

## About ApprenticeBench

- Year: 2026
- Tasks: 100 vendor bills processed in sequence inside a simulated construction company
- Format: Cumulative success rate over 100 bills
- Difficulty: Long-horizon computer use with offline and online continual learning
- Paper: [ApprenticeBench: a step change in AI's job readiness](https://neocognition.io/blog/apprentice-bench/)

NeoCognition simulates Acme Home Builders, a California construction company with a full year of business. The agent joins in May 2026 with six months of historical bills, the company handbook, and an Odoo ERP tutorial, then processes 100 incoming bills in order. It receives immediate feedback from the accounts-payable manager during the first month and only sparse month-end feedback afterward. Bills embed unit, price, tax, cost-code, vendor-documentation, and approval-policy challenges that two construction accounting professionals validated as realistic and learnable. Agents run inside coding-agent harnesses (Codex for OpenAI models, Claude Code for everyone else) extended with computer-use tools, a 500K context, and persistent Markdown memory notes. We mirror every published GUI and API run as display-only evidence.

ApprenticeBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (48 setting-level runs)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Fable 5.1](/models/claude-fable-5-1) | Claude Code · max reasoning · GUI (computer use) · $18.23 per task | Anthropic | 72% |
| 2 | [Claude Fable 5.1](/models/claude-fable-5-1) | Claude Code · max reasoning · API (MCP tools) · $6.95 per task | Anthropic | 70% |
| 3 | [GPT-6 Astra](/models/gpt-6-astra) | Codex · max reasoning · GUI (computer use) · $20.51 per task | OpenAI | 68% |
| 4 | [GPT-6 Astra](/models/gpt-6-astra) | Codex · max reasoning · API (MCP tools) · $9.62 per task | OpenAI | 65% |
| 5 | [GPT-6 Astra](/models/gpt-6-astra) | Codex · xhigh reasoning · GUI (computer use) · $14.37 per task | OpenAI | 61% |
| 6 | [Claude Fable 5.1](/models/claude-fable-5-1) | Claude Code · xhigh reasoning · GUI (computer use) · $13.87 per task | Anthropic | 59% |
| 7 | [Claude Opus 5](/models/claude-opus-5) | Claude Code · max reasoning · API (MCP tools) · $5.28 per task | Anthropic | 49% |
| 8 | [Grok 4.6](/models/grok-4-6) | Claude Code · max reasoning · API (MCP tools) · $1.44 per task | xAI | 45% |
| 9 | [Claude Fable 5](/models/claude-fable) | Claude Code · max reasoning · API (MCP tools) · $7.33 per task | Anthropic | 45% |
| 10 | [Claude Fable 5](/models/claude-fable) | Claude Code · xhigh reasoning · GUI (computer use) · $33.10 per task | Anthropic | 43% |
| 11 | [Grok 4.5](/models/grok-4-5) | Claude Code · max reasoning · API (MCP tools) · $0.83 per task | xAI | 38% |
| 12 | [Claude Opus 5](/models/claude-opus-5) | Claude Code · max reasoning · GUI (computer use) · $20.07 per task | Anthropic | 36% |
| 13 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Claude Code · max reasoning · API (MCP tools) · $1.13 per task | Google | 34% |
| 14 | [Claude Fable 5](/models/claude-fable) | Claude Code · max reasoning · GUI (computer use) · $28.68 per task | Anthropic | 34% |
| 15 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Claude Code · max reasoning · API (MCP tools) · $0.61 per task | Google | 32% |
| 16 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | Codex · max reasoning · API (MCP tools) · $2.04 per task | OpenAI | 30% |
| 17 | [Claude Opus 4.8](/models/claude-opus-4-8) | Claude Code · max reasoning · API (MCP tools) · $4.59 per task | Anthropic | 28% |
| 18 | [GPT-5.5](/models/gpt-5-5) | Codex · max reasoning · API (MCP tools) · $2.82 per task | OpenAI | 27% |
| 19 | [Claude Sonnet 5](/models/claude-sonnet-5) | Claude Code · max reasoning · API (MCP tools) · $1.91 per task | Anthropic | 26% |
| 20 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | Codex · max reasoning · GUI (computer use) · $8.68 per task | OpenAI | 26% |
| 21 | [Kimi K3](/models/kimi-k3) | Claude Code · max reasoning · API (MCP tools) · $1.59 per task | Moonshot AI | 25% |
| 22 | [Muse Spark 1.3](/models/muse-spark-1-3) | Claude Code · max reasoning · API (MCP tools) · $0.57 per task | Meta | 24% |
| 23 | [Qwen3.8 Max](/models/qwen3-8-max) | Claude Code · max reasoning · API (MCP tools) · $1.39 per task | Alibaba | 24% |
| 24 | [Gemini 3.8 Flash](/models/gemini-3-8-flash) | Claude Code · max reasoning · GUI (computer use) · $4.10 per task | Google | 24% |
| 25 | [GLM-5.3-Flash](/models/glm-5-3-flash) | Claude Code · max reasoning · API (MCP tools) · $0.07 per task | Z.AI | 21% |
| 26 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | Claude Code · max reasoning · API (MCP tools) · $0.46 per task | DeepSeek | 21% |
| 27 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | Codex · max reasoning · API (MCP tools) · $1.13 per task | OpenAI | 21% |
| 28 | [GPT-5.5](/models/gpt-5-5) | Codex · max reasoning · GUI (computer use) · $12.68 per task | OpenAI | 20% |
| 29 | [Muse Spark 1.2](/models/muse-spark-1-2) | Claude Code · max reasoning · API (MCP tools) · $2.38 per task | Meta | 19% |
| 30 | [Muse Spark 1.3](/models/muse-spark-1-3) | Claude Code · max reasoning · GUI (computer use) · $7.93 per task | Meta | 19% |
| 31 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | Claude Code · max reasoning · API (MCP tools) · $0.07 per task | DeepSeek | 18% |
| 32 | [Kimi K3](/models/kimi-k3) | Claude Code · max reasoning · GUI (computer use) · $25.83 per task | Moonshot AI | 18% |
| 33 | [GLM-5.2](/models/glm-5-2) | Claude Code · max reasoning · API (MCP tools) · $0.60 per task | Z.AI | 17% |
| 34 | [Gemini 3.7 Flash](/models/gemini-3-7-flash) | Claude Code · max reasoning · GUI (computer use) · $3.92 per task | Google | 16% |
| 35 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | Codex · max reasoning · GUI (computer use) · $6.15 per task | OpenAI | 16% |
| 36 | [Claude Sonnet 5](/models/claude-sonnet-5) | Claude Code · max reasoning · GUI (computer use) · $14.62 per task | Anthropic | 16% |
| 37 | [Claude Opus 4.7](/models/claude-opus-4-7) | Claude Code · max reasoning · API (MCP tools) · $2.42 per task | Anthropic | 14% |
| 38 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | Codex · max reasoning · API (MCP tools) · $0.16 per task | OpenAI | 13% |
| 39 | [Grok 4.6](/models/grok-4-6) | Claude Code · max reasoning · GUI (computer use) · $35.79 per task | xAI | 13% |
| 40 | [Gemini 3.6 Flash](/models/gemini-3-6-flash) | Claude Code · max reasoning · API (MCP tools) · $0.46 per task | Google | 12% |
| 41 | [Muse Spark 1.1](/models/muse-spark-1-1) | Claude Code · max reasoning · API (MCP tools) · $0.90 per task | Meta | 12% |
| 42 | [GPT-5.4](/models/gpt-5-4) | Codex · max reasoning · GUI (computer use) · $5.95 per task | OpenAI | 11% |
| 43 | [Qwen 3.8 Flash · max · API](https://neocognition.io/blog/apprentice-bench/#the-diminishing-cua-tax-gui-vs-api) | Claude Code · max reasoning · API (MCP tools) · $0.16 per task | Alibaba | 10% |
| 44 | [Kimi K2.6](/models/kimi-2-6) | Claude Code · max reasoning · API (MCP tools) · $0.69 per task | Moonshot AI | 10% |
| 45 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | Codex · max reasoning · GUI (computer use) · $0.77 per task | OpenAI | 7% |
| 46 | [Claude Opus 4.7](/models/claude-opus-4-7) | Claude Code · max reasoning · GUI (computer use) · $26.04 per task | Anthropic | 7% |
| 47 | [Claude Opus 4.6](/models/claude-opus-4-6) | Claude Code · max reasoning · GUI (computer use) · $23.68 per task | Anthropic | 5% |
| 48 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Claude Code · max reasoning · GUI (computer use) · $14.68 per task | Anthropic | 2% |

## FAQ

### What does ApprenticeBench measure?

ApprenticeBench measures whether an AI agent can learn and do a real knowledge job end to end. The first instantiation is an accounts-payable clerk role at a simulated construction company. The agent must read invoices, cross-check purchase orders and policies, message vendors for corrections, and post bills in the Odoo ERP system, learning company conventions from historical data and mentor feedback across 100 bills.

### Which model leads the published ApprenticeBench results?

Claude Fable 5.1 at max effort leads the September 11, 2026 GUI results with a 72% cumulative success rate at $18.23 per task. GPT-6 Astra follows at 68%. The best human tester passed 51% of the same bills at an estimated $7.21 per task, so the two leading agents are more accurate but more expensive than the human.

### Why is ApprenticeBench display only?

Each result combines the model with a reasoning effort, a coding-agent harness, a GUI or API interface setting, a memory-note learning strategy, and a single non-resettable pass through the bills. The snapshot is useful for comparing those exact setups, but it does not enter BenchLM's model-only rankings.
