Market-Bench
We show this table for reference; we do not rank on it.
A quantitative-trading implementation benchmark that asks models to build backtesters under market-book liquidity and execution-delay constraints, then compares their outputs with a verifier.
Mean MAE on Market-Bench — September 23, 2026 snapshot
We mirror the published mean mae view for Market-Bench. Grok 4 leads the public snapshot at 443.24, followed by GPT-5.2 (969.39) and Gemini 3 Pro (1744.27). We do not use these results to rank models overall.
Grok 4
xAI
GPT-5.2
OpenAI
Gemini 3 Pro
13 modelsAgenticCurrentDisplay onlyUpdated September 23, 2026 snapshot
Mean MAE table (13 models)
ScoreHow we show Market-Bench
We mirror AfterQuery's Market-Bench table captured on September 23, 2026 snapshot. The benchmark asks models to implement backtesters for 3 quantitative-trading strategies, then measures mean absolute error against verifiable reference outputs.
Mean MAE is lower-is-better. The strategies include market-book liquidity and execution-delay constraints, so this is a code-and-simulation result rather than a prediction of investment returns. We keep it display-only and do not mix its unbounded error scale with percentage benchmarks.
Snapshot
The published Market-Bench snapshot places Grok 4 first at 443.24. The third row is 1301.03 score units higher. The broader top-10 range is 9230.97 score units, so the table still separates the published systems.
13 models have been evaluated on Market-Bench. The benchmark falls in the Agentic category. Market-Bench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About Market-Bench
Year
2025
Tasks
3 quantitative-trading strategies
Format
Backtester implementation scored by mean absolute error
Difficulty
Market simulation and quantitative coding
Market-Bench covers scheduled single-stock execution, pairs mean reversion, and dynamic delta hedging. The public table reports mean absolute error across the strategies; lower values are better. The unbounded error scale stays display-only and is not mixed with percentage benchmarks.
Freshness and provenance
Version
Market-Bench 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does Market-Bench measure?
A quantitative-trading implementation benchmark that asks models to build backtesters under market-book liquidity and execution-delay constraints, then compares their outputs with a verifier.
Which model leads the published Market-Bench snapshot?
Grok 4 currently leads the published Market-Bench snapshot with 443.24 mean mae. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on Market-Bench?
The September 23, 2026 snapshot snapshot contains 13 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.