ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job (ApprenticeBench)
We mirror this table; we do not rank on it.
Tests whether a computer-use agent can learn a real accounts-payable job on the job, processing 100 vendor bills in a company ERP system with only the handbook, historical records, and mentor feedback a new hire would get.
Success rate on ApprenticeBench — September 18, 2026 snapshot
We mirror the published success rate view for ApprenticeBench. Claude Fable 5.1 leads the public snapshot at 72%, followed by Claude Fable 5.1 (70%) and GPT-6 Astra (68%). We do not use these results to rank models overall.
Claude Fable 5.1
Anthropic
Claude Code · max reasoning
GUI (computer use) · $18.23 per task
Claude Fable 5.1
Anthropic
Claude Code · max reasoning
API (MCP tools) · $6.95 per task
GPT-6 Astra
OpenAI
Codex · max reasoning
GUI (computer use) · $20.51 per task
48 setting-level runsAgenticCurrentDisplay onlyUpdated September 18, 2026 snapshot
Success rate table (48 setting-level runs)
ScoreHow to read this leaderboard
A bill counts as solved only when every field, correction, vendor exchange, and approval matches the ground truth. The GUI setting is the benchmark's main setting; the API setting swaps the ERP interface for MCP tools with feature parity, and the gap between them is the CUA tax.
Operator receipt: 48 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5.1 at 72%.
Honest limit: Every score belongs to a full setup: model, reasoning effort, coding-agent harness, interface setting, memory strategy, and one non-resettable pass through the bills. Only one job (construction accounts payable) is instantiated so far, and the benchmark tasks are private. Costs are per-task API spend reported by the benchmark owner.
How we show ApprenticeBench
NeoCognition publishes ApprenticeBench results as charts rather than a table, so we mirror every run's published label from the September 18, 2026 snapshot: 21 GUI (computer-use) runs across 18 models and 27 API (MCP-tool) runs across 27 models. The score is the cumulative success rate over 100 accounts-payable bills processed in order.
The GUI setting is the benchmark's main setting. The best human tester passed 51% of the same bills at an estimated $7.21 per task. Runs are labeled with their interface setting and per-task cost; unlabeled source runs use max reasoning effort, and the xhigh runs the source charts single out stay separate.
Agents run inside Codex (OpenAI models) or Claude Code (all others) extended with computer-use tools, a 500K context, and persistent Markdown memory notes. Each score belongs to that full setup and a single non-resettable pass, so the table is display only and does not enter model rankings.
Snapshot
The published ApprenticeBench snapshot places Claude Fable 5.1 first at 72%. The third row is 4 points behind. The broader top-10 range is 29 points, so the table still separates the published systems.
48 setting-level runs have been evaluated on ApprenticeBench. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. ApprenticeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About ApprenticeBench
Year
2026
Tasks
100 vendor bills processed in sequence inside a simulated construction company
Format
Cumulative success rate over 100 bills
Difficulty
Long-horizon computer use with offline and online continual learning
NeoCognition simulates Acme Home Builders, a California construction company with a full year of business. The agent joins in May 2026 with six months of historical bills, the company handbook, and an Odoo ERP tutorial, then processes 100 incoming bills in order. It receives immediate feedback from the accounts-payable manager during the first month and only sparse month-end feedback afterward. Bills embed unit, price, tax, cost-code, vendor-documentation, and approval-policy challenges that two construction accounting professionals validated as realistic and learnable. Agents run inside coding-agent harnesses (Codex for OpenAI models, Claude Code for everyone else) extended with computer-use tools, a 500K context, and persistent Markdown memory notes. We mirror every published GUI and API run as display-only evidence.
Freshness and provenance
Version
ApprenticeBench (accounts-payable instantiation)
Refresh cadence
Quarterly
Staleness state
Current
Question availability
100 private bills in a simulated company; results published on the NeoCognition blog
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does ApprenticeBench measure?
ApprenticeBench measures whether an AI agent can learn and do a real knowledge job end to end. The first instantiation is an accounts-payable clerk role at a simulated construction company. The agent must read invoices, cross-check purchase orders and policies, message vendors for corrections, and post bills in the Odoo ERP system, learning company conventions from historical data and mentor feedback across 100 bills.
Which model leads the published ApprenticeBench results?
Claude Fable 5.1 at max effort leads the September 11, 2026 GUI results with a 72% cumulative success rate at $18.23 per task. GPT-6 Astra follows at 68%. The best human tester passed 51% of the same bills at an estimated $7.21 per task, so the two leading agents are more accurate but more expensive than the human.
Why is ApprenticeBench display only?
Each result combines the model with a reasoning effort, a coding-agent harness, a GUI or API interface setting, a memory-note learning strategy, and a single non-resettable pass through the bills. The snapshot is useful for comparing those exact setups, but it does not enter BenchLM's model-only rankings.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.