# Market-Bench

> A quantitative-trading implementation benchmark that asks models to build backtesters under market-book liquidity and execution-delay constraints, then compares their outputs with a verifier.

Canonical page: https://benchlm.ai/benchmarks/marketbench

- Category: [Agentic](/agentic)
- Last updated: September 23, 2026 snapshot

## About Market-Bench

- Year: 2025
- Tasks: 3 quantitative-trading strategies
- Format: Backtester implementation scored by mean absolute error
- Difficulty: Market simulation and quantitative coding
- Paper: [Market-Bench](https://arxiv.org/abs/2512.12264)

Market-Bench covers scheduled single-stock execution, pairs mean reversion, and dynamic delta hedging. The public table reports mean absolute error across the strategies; lower values are better. The unbounded error scale stays display-only and is not mixed with percentage benchmarks.

Market-Bench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (13 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Grok 4](/models/grok-4) | xAI | 443.24 |
| 2 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 969.39 |
| 3 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 1744.27 |
| 4 | [GPT-5.1-Codex-Max](/models/gpt-5-1-codex-max) | OpenAI | 4242.93 |
| 5 | [DeepSeek V3.2](/models/deepseek-v3-2) | DeepSeek | 4575.64 |
| 6 | [Claude Sonnet 4.5](/models/claude-sonnet-4-5) | Anthropic | 5126.76 |
| 7 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 6039.62 |
| 8 | [Command A](https://www.afterquery.com/leaderboard/market-bench) | Cohere | 6562.11 |
| 9 | [Nova Premier](https://www.afterquery.com/leaderboard/market-bench) | Amazon | 7740.26 |
| 10 | [Llama 3.1 Nemotron Ultra](https://www.afterquery.com/leaderboard/market-bench) | NVIDIA | 9674.21 |
| 11 | [Llama 4 Maverick](/models/llama-4-maverick) | Meta | 10202.18 |
| 12 | [Mistral Large](https://www.afterquery.com/leaderboard/market-bench) | Mistral | 30605.53 |
| 13 | [Qwen3 Max](/models/qwen3-max) | Alibaba | 159144490.12 |

## FAQ

### What does Market-Bench measure?

A quantitative-trading implementation benchmark that asks models to build backtesters under market-book liquidity and execution-delay constraints, then compares their outputs with a verifier.

### Which model leads the published Market-Bench snapshot?

Grok 4 currently leads the published Market-Bench snapshot with a score of 443.24.

### How many models are evaluated on Market-Bench?

The September 23, 2026 snapshot contains 13 AI models.

### Does Market-Bench affect BenchLM's overall score?

Not directly. Market-Bench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
