Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job (ApprenticeBench)

We mirror this table; we do not rank on it.

Tests whether a computer-use agent can learn a real accounts-payable job on the job, processing 100 vendor bills in a company ERP system with only the handbook, historical records, and mentor feedback a new hire would get.

Success rate on ApprenticeBench — September 18, 2026 snapshot

We mirror the published success rate view for ApprenticeBench. Claude Fable 5.1 leads the public snapshot at 72%, followed by Claude Fable 5.1 (70%) and GPT-6 Astra (68%). We do not use these results to rank models overall.

48 setting-level runsAgenticCurrentDisplay onlyUpdated September 18, 2026 snapshot

Success rate table (48 setting-level runs)

Score
1
Claude Fable 5.1Anthropic · ClosedClaude Code · max reasoningGUI (computer use) · $18.23 per task
72%
2
Claude Fable 5.1Anthropic · ClosedClaude Code · max reasoningAPI (MCP tools) · $6.95 per task
70%
3
GPT-6 AstraOpenAI · ClosedCodex · max reasoningGUI (computer use) · $20.51 per task
68%
4
GPT-6 AstraOpenAI · ClosedCodex · max reasoningAPI (MCP tools) · $9.62 per task
65%
5
GPT-6 AstraOpenAI · ClosedCodex · xhigh reasoningGUI (computer use) · $14.37 per task
61%
6
Claude Fable 5.1Anthropic · ClosedClaude Code · xhigh reasoningGUI (computer use) · $13.87 per task
59%
7
Claude Opus 5Anthropic · ClosedClaude Code · max reasoningAPI (MCP tools) · $5.28 per task
49%
8
Grok 4.6xAI · ClosedClaude Code · max reasoningAPI (MCP tools) · $1.44 per task
45%
9
Claude Fable 5Anthropic · ClosedClaude Code · max reasoningAPI (MCP tools) · $7.33 per task
45%
10
Claude Fable 5Anthropic · ClosedClaude Code · xhigh reasoningGUI (computer use) · $33.10 per task
43%
11
Grok 4.5xAI · ClosedClaude Code · max reasoningAPI (MCP tools) · $0.83 per task
38%
12
Claude Opus 5Anthropic · ClosedClaude Code · max reasoningGUI (computer use) · $20.07 per task
36%
13
Gemini 3.8 FlashGoogle · ClosedClaude Code · max reasoningAPI (MCP tools) · $1.13 per task
34%
14
Claude Fable 5Anthropic · ClosedClaude Code · max reasoningGUI (computer use) · $28.68 per task
34%
15
Gemini 3.7 FlashGoogle · ClosedClaude Code · max reasoningAPI (MCP tools) · $0.61 per task
32%
16
GPT-5.6 SolOpenAI · ClosedCodex · max reasoningAPI (MCP tools) · $2.04 per task
30%
17
Claude Opus 4.8Anthropic · ClosedClaude Code · max reasoningAPI (MCP tools) · $4.59 per task
28%
18
GPT-5.5OpenAI · ClosedCodex · max reasoningAPI (MCP tools) · $2.82 per task
27%
19
Claude Sonnet 5Anthropic · ClosedClaude Code · max reasoningAPI (MCP tools) · $1.91 per task
26%
20
GPT-5.6 SolOpenAI · ClosedCodex · max reasoningGUI (computer use) · $8.68 per task
26%
21
Kimi K3Moonshot AI · ClosedClaude Code · max reasoningAPI (MCP tools) · $1.59 per task
25%
22
Muse Spark 1.3Meta · ClosedClaude Code · max reasoningAPI (MCP tools) · $0.57 per task
24%
23
Qwen3.8 MaxAlibaba · Open weightClaude Code · max reasoningAPI (MCP tools) · $1.39 per task
24%
24
Gemini 3.8 FlashGoogle · ClosedClaude Code · max reasoningGUI (computer use) · $4.10 per task
24%
25
GLM-5.3-FlashZ.AI · Open weightClaude Code · max reasoningAPI (MCP tools) · $0.07 per task
21%
26
DeepSeek V4 Pro 0813DeepSeek · ClosedClaude Code · max reasoningAPI (MCP tools) · $0.46 per task
21%
27
GPT-5.6 TerraOpenAI · ClosedCodex · max reasoningAPI (MCP tools) · $1.13 per task
21%
28
GPT-5.5OpenAI · ClosedCodex · max reasoningGUI (computer use) · $12.68 per task
20%
29
Muse Spark 1.2Meta · ClosedClaude Code · max reasoningAPI (MCP tools) · $2.38 per task
19%
30
Muse Spark 1.3Meta · ClosedClaude Code · max reasoningGUI (computer use) · $7.93 per task
19%
31
DeepSeek V4 Flash 0731DeepSeek · ClosedClaude Code · max reasoningAPI (MCP tools) · $0.07 per task
18%
32
Kimi K3Moonshot AI · ClosedClaude Code · max reasoningGUI (computer use) · $25.83 per task
18%
33
GLM-5.2Z.AI · Open weightClaude Code · max reasoningAPI (MCP tools) · $0.60 per task
17%
34
Gemini 3.7 FlashGoogle · ClosedClaude Code · max reasoningGUI (computer use) · $3.92 per task
16%
35
GPT-5.6 TerraOpenAI · ClosedCodex · max reasoningGUI (computer use) · $6.15 per task
16%
36
Claude Sonnet 5Anthropic · ClosedClaude Code · max reasoningGUI (computer use) · $14.62 per task
16%
37
Claude Opus 4.7Anthropic · ClosedClaude Code · max reasoningAPI (MCP tools) · $2.42 per task
14%
38
GPT-5.6 LunaOpenAI · ClosedCodex · max reasoningAPI (MCP tools) · $0.16 per task
13%
39
Grok 4.6xAI · ClosedClaude Code · max reasoningGUI (computer use) · $35.79 per task
13%
40
Gemini 3.6 FlashGoogle · ClosedClaude Code · max reasoningAPI (MCP tools) · $0.46 per task
12%
41
Muse Spark 1.1Meta · ClosedClaude Code · max reasoningAPI (MCP tools) · $0.90 per task
12%
42
GPT-5.4OpenAI · ClosedCodex · max reasoningGUI (computer use) · $5.95 per task
11%
43
Qwen 3.8 Flash · max · APIAlibabaClaude Code · max reasoningAPI (MCP tools) · $0.16 per task
10%
44
Kimi K2.6Moonshot AI · Open weightClaude Code · max reasoningAPI (MCP tools) · $0.69 per task
10%
45
GPT-5.6 LunaOpenAI · ClosedCodex · max reasoningGUI (computer use) · $0.77 per task
7%
46
Claude Opus 4.7Anthropic · ClosedClaude Code · max reasoningGUI (computer use) · $26.04 per task
7%
47
Claude Opus 4.6Anthropic · ClosedClaude Code · max reasoningGUI (computer use) · $23.68 per task
5%
48
Claude Sonnet 4.6Anthropic · ClosedClaude Code · max reasoningGUI (computer use) · $14.68 per task
2%

How to read this leaderboard

A bill counts as solved only when every field, correction, vendor exchange, and approval matches the ground truth. The GUI setting is the benchmark's main setting; the API setting swaps the ERP interface for MCP tools with feature parity, and the gap between them is the CUA tax.

Operator receipt: 48 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5.1 at 72%.

Honest limit: Every score belongs to a full setup: model, reasoning effort, coding-agent harness, interface setting, memory strategy, and one non-resettable pass through the bills. Only one job (construction accounts payable) is instantiated so far, and the benchmark tasks are private. Costs are per-task API spend reported by the benchmark owner.

How we show ApprenticeBench

NeoCognition publishes ApprenticeBench results as charts rather than a table, so we mirror every run's published label from the September 18, 2026 snapshot: 21 GUI (computer-use) runs across 18 models and 27 API (MCP-tool) runs across 27 models. The score is the cumulative success rate over 100 accounts-payable bills processed in order.

The GUI setting is the benchmark's main setting. The best human tester passed 51% of the same bills at an estimated $7.21 per task. Runs are labeled with their interface setting and per-task cost; unlabeled source runs use max reasoning effort, and the xhigh runs the source charts single out stay separate.

Agents run inside Codex (OpenAI models) or Claude Code (all others) extended with computer-use tools, a 500K context, and persistent Markdown memory notes. Each score belongs to that full setup and a single non-resettable pass, so the table is display only and does not enter model rankings.

The published ApprenticeBench snapshot places Claude Fable 5.1 first at 72%. The third row is 4 points behind. The broader top-10 range is 29 points, so the table still separates the published systems.

48 setting-level runs have been evaluated on ApprenticeBench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. ApprenticeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About ApprenticeBench

Year

2026

Tasks

100 vendor bills processed in sequence inside a simulated construction company

Format

Cumulative success rate over 100 bills

Difficulty

Long-horizon computer use with offline and online continual learning

NeoCognition simulates Acme Home Builders, a California construction company with a full year of business. The agent joins in May 2026 with six months of historical bills, the company handbook, and an Odoo ERP tutorial, then processes 100 incoming bills in order. It receives immediate feedback from the accounts-payable manager during the first month and only sparse month-end feedback afterward. Bills embed unit, price, tax, cost-code, vendor-documentation, and approval-policy challenges that two construction accounting professionals validated as realistic and learnable. Agents run inside coding-agent harnesses (Codex for OpenAI models, Claude Code for everyone else) extended with computer-use tools, a 500K context, and persistent Markdown memory notes. We mirror every published GUI and API run as display-only evidence.

Freshness and provenance

Version

ApprenticeBench (accounts-payable instantiation)

Refresh cadence

Quarterly

Staleness state

Current

Question availability

100 private bills in a simulated company; results published on the NeoCognition blog

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does ApprenticeBench measure?

ApprenticeBench measures whether an AI agent can learn and do a real knowledge job end to end. The first instantiation is an accounts-payable clerk role at a simulated construction company. The agent must read invoices, cross-check purchase orders and policies, message vendors for corrections, and post bills in the Odoo ERP system, learning company conventions from historical data and mentor feedback across 100 bills.

Which model leads the published ApprenticeBench results?

Claude Fable 5.1 at max effort leads the September 11, 2026 GUI results with a 72% cumulative success rate at $18.23 per task. GPT-6 Astra follows at 68%. The best human tester passed 51% of the same bills at an estimated $7.21 per task, so the two leading agents are more accurate but more expensive than the human.

Why is ApprenticeBench display only?

Each result combines the model with a reasoning effort, a coding-agent harness, a GUI or API interface setting, a memory-note learning strategy, and a single non-resettable pass through the bills. The snapshot is useful for comparing those exact setups, but it does not enter BenchLM's model-only rankings.

Last updated: September 18, 2026 snapshot · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.