Skip to main content
BenchLM
Recommendation

Best Value Agentic AI Model in 2026 — Cost-Adjusted Rankings

As of October 1, 2026, the top model in best value agentic ai model on the BenchLM leaderboard is MiMo-V2.6-Flash with a score of 195.3.

Bottom line: Agentic tasks are token-intensive — value matters more here than almost anywhere.

Ranking data as of

Full Rankings (76 models)

1
MiMo-V2.6-Flash
Xiaomi·Open Weight·1M

195.29

Score/$

Score: 54.7 · $0.28/1M

2
GPT-6 Luna
OpenAI·Proprietary·1.05M

106.06

Score/$

Score: 53 · $0.5/1M

3
MiMo-V2.6-Pro
Xiaomi·Open Weight·1M

76.37

Score/$

Score: 66.4 · $0.87/1M

4
Mercury 2.5
Inception·Proprietary·260K

75.13

Score/$

Score: 11.3 · $0.15/1M

5
DeepSeek V4.1 Flash
DeepSeek·Open Weight·1M

50.89

Score/$

Score: 61.1 · $1.2/1M

6
Ministral 3 3B
Mistral·Open Weight·128K

47.9

Score/$

Score: 4.8 · $0.1/1M

7
GPT-5.6 Luna
OpenAI·Proprietary·1.05M

45.88

Score/$

Score: 55.1 · $1.2/1M

8
Ministral 3 8B
Mistral·Open Weight·128K

36.53

Score/$

Score: 5.5 · $0.15/1M

9
MiniMax M3
MiniMax·Open Weight·1M

33.23

Score/$

Score: 39.9 · $1.2/1M

10
Ministral 3 14B
Mistral·Open Weight·128K

31.2

Score/$

Score: 6.2 · $0.2/1M

11
Step 3.7 Flash
StepFun·Open Weight·256K

31.14

Score/$

Score: 35.8 · $1.15/1M

12
GPT-5.4 nano
OpenAI·Proprietary·400K

25.45

Score/$

Score: 31.8 · $1.25/1M

13
Inkling-Small
Thinking Machines Lab·Open Weight·1M

24.53

Score/$

Score: 35.3 · $1.44/1M

14
MiniMax M2.7
MiniMax·Open Weight·200K

24.29

Score/$

Score: 29.2 · $1.2/1M

15
MiniMax M2.5
MiniMax·Proprietary·128K

23.81

Score/$

Score: 28.6 · $1.2/1M

16
Step 5 Preview
StepFun·Pending·1M

23.51

Score/$

Score: 63.5 · $2.7/1M

17
Mistral Small 4
Mistral·Open Weight·256K

22.12

Score/$

Score: 13.3 · $0.6/1M

18
Quasar 438B
Multiverse Computing·Proprietary·1M

21.62

Score/$

Score: 38.9 · $1.8/1M

19
Mercury 2
Inception·Proprietary·128K

19.8

Score/$

Score: 14.9 · $0.75/1M

20
Gemini 3.1 Flash-Lite
Google·Proprietary·1M

18.48

Score/$

Score: 27.7 · $1.5/1M

21
Gemini 3.8 Flash
Google·Proprietary·1M

17.41

Score/$

Score: 65.3 · $3.75/1M

22
Gemini 3.7 Flash
Google·Proprietary·1M

15.47

Score/$

Score: 58 · $3.75/1M

23
Trinity-Large-Thinking
Arcee AI·Open Weight·512K

14.09

Score/$

Score: 12.7 · $0.9/1M

24
Muse Spark 1.2
Meta·Proprietary·1M

13.96

Score/$

Score: 59.4 · $4.25/1M

25
Grok Build 0.1
xAI·Proprietary·256K

13.93

Score/$

Score: 27.9 · $2/1M

26
GPT-5 mini
OpenAI·Proprietary·128K

13.85

Score/$

Score: 27.7 · $2/1M

27
Gemini 3.5 Flash-Lite
Google·Proprietary·1M

13.65

Score/$

Score: 34.1 · $2.5/1M

28
DeepSeek V4 Pro 0813
DeepSeek·Open Weight·1M

13.38

Score/$

Score: 53 · $3.96/1M

29
GLM-5.2
Z.AI·Open Weight·1M

12.75

Score/$

Score: 56.1 · $4.4/1M

30
Kimi K2.5
Moonshot AI·Open Weight·256K

12.65

Score/$

Score: 38 · $3/1M

31
GLM-5
Z.AI·Open Weight·200K

12.13

Score/$

Score: 38.8 · $3.2/1M

32
Gemini 3.6 Flash
Google·Proprietary·1M

11.99

Score/$

Score: 45 · $3.75/1M

33
Grok 4.6
xAI·Proprietary·500K

11.35

Score/$

Score: 68.1 · $6/1M

34
Kimi K2.6
Moonshot AI·Open Weight·256K

10.84

Score/$

Score: 43.4 · $4/1M

35
Gemini 3 Flash
Google·Proprietary·1M

10.57

Score/$

Score: 31.7 · $3/1M

36
Grok 4.3
xAI·Proprietary·1M

10.4

Score/$

Score: 26 · $2.5/1M

37
GLM-5.1
Z.AI·Open Weight·203K

9.61

Score/$

Score: 42.3 · $4.4/1M

38
Grok 4.20
xAI·Proprietary·2M

9.6

Score/$

Score: 24 · $2.5/1M

39
Grok 4.5
xAI·Proprietary·500K

9.42

Score/$

Score: 56.5 · $6/1M

40
Kimi K2.7 Code
Moonshot AI·Open Weight·256K

9.4

Score/$

Score: 37.6 · $4/1M

41
DeepSeek V3
DeepSeek·Open Weight·128K

9.38

Score/$

Score: 10.3 · $1.1/1M

42
Celeris-1
Celeris·Proprietary·128K

9.24

Score/$

Score: 6.5 · $0.7/1M

43
Inkling
Thinking Machines Lab·Open Weight·1M

7.73

Score/$

Score: 36.2 · $4.68/1M

44
GPT-5.4 mini
OpenAI·Proprietary·400K

7.57

Score/$

Score: 34.1 · $4.5/1M

45
Gemini 4 Argon
Google·Proprietary·

7.42

Score/$

Score: 74.2 · $10/1M

46
Mistral Large 3
Mistral·Proprietary·128K

6.95

Score/$

Score: 10.4 · $1.5/1M

47
Claude Sonnet 5.5
Anthropic·Proprietary·1M

6.8

Score/$

Score: 68 · $10/1M

48
GPT-6.1 Sol
OpenAI·Proprietary·1.05M

6.78

Score/$

Score: 67.8 · $10/1M

49
Claude Sonnet 5
Anthropic·Proprietary·1M

6.46

Score/$

Score: 64.6 · $10/1M

50
GPT-6 Sol
OpenAI·Proprietary·1.05M

5.91

Score/$

Score: 59.1 · $10/1M

51
Gemini 3.5 Flash
Google·Proprietary·1M

5.6

Score/$

Score: 50.4 · $9/1M

52
GPT-5.6 Terra
OpenAI·Proprietary·1.05M

4.94

Score/$

Score: 59.2 · $12/1M

53
Kimi K3
Moonshot AI·Pending·1.05M

4.52

Score/$

Score: 67.8 · $15/1M

54
Claude Opus 5.5
Anthropic·Proprietary·1M

4.4

Score/$

Score: 88.1 · $20/1M

55
Claude Haiku 4.5
Anthropic·Proprietary·200K

4.36

Score/$

Score: 21.8 · $5/1M

56
GPT-5.3 Codex
OpenAI·Proprietary·400K

3.96

Score/$

Score: 55.4 · $14/1M

57
GPT-5.6 Sol
OpenAI·Proprietary·1.05M

3.4

Score/$

Score: 68 · $20/1M

58
GPT-5.4
OpenAI·Proprietary·1.05M

3.3

Score/$

Score: 49.6 · $15/1M

59
Gemini 3.1 Pro
Google·Proprietary·1M

3.2

Score/$

Score: 38.4 · $12/1M

60
Claude Opus 5
Anthropic·Proprietary·

3.1

Score/$

Score: 77.6 · $25/1M

61
GPT-5.2
OpenAI·Proprietary·400K

3.06

Score/$

Score: 42.8 · $14/1M

62
Claude Sonnet 4.6
Anthropic·Proprietary·200K

2.9

Score/$

Score: 43.4 · $15/1M

63
GPT-5.2-Codex
OpenAI·Proprietary·400K

2.88

Score/$

Score: 40.3 · $14/1M

64
Mistral Medium 3.5 128B
Mistral·Open Weight·256K

2.57

Score/$

Score: 19.3 · $7.5/1M

65
Claude Opus 4.8
Anthropic·Proprietary·1M

2.46

Score/$

Score: 61.6 · $25/1M

66
Gemini 2.5 Pro
Google·Proprietary·1M

2.44

Score/$

Score: 24.4 · $10/1M

67
Claude Opus 4.7 (Adaptive)
Anthropic·Proprietary·1M

2.37

Score/$

Score: 59.2 · $25/1M

68
Claude Sonnet 4.5
Anthropic·Proprietary·200K

2.26

Score/$

Score: 33.9 · $15/1M

69
Claude Opus 4.7
Anthropic·Proprietary·1M

2.13

Score/$

Score: 53.2 · $25/1M

70
GPT-5.5
OpenAI·Proprietary·1M

1.98

Score/$

Score: 59.5 · $30/1M

71
Claude Opus 4.6
Anthropic·Proprietary·1M

1.78

Score/$

Score: 44.4 · $25/1M

72
Command A+
Cohere·Open Weight·128K

1.59

Score/$

Score: 15.9 · $10/1M

73
Claude Fable 5.1
Anthropic·Proprietary·1M

1.58

Score/$

Score: 78.9 · $50/1M

74
Claude Fable 5
Anthropic·Proprietary·1M+

1.48

Score/$

Score: 74 · $50/1M

75
GPT-6 Astra
OpenAI·Proprietary·1.05M

1.41

Score/$

Score: 70.7 · $50/1M

76
Claude Opus 4.5
Anthropic·Proprietary·200K

1.24

Score/$

Score: 31.1 · $25/1M

How to choose

Key Takeaways

The best value model is MiMo-V2.6-Flash by Xiaomi with a provisional Score/$ ratio of 195.29 (score: 54.7, output: $0.28/1M tokens).

The best open-weight model is MiMo-V2.6-Flash at position #1.

76 models are included in this ranking.

Score in Context

What these scores mean

Value scores divide the weighted agentic score by output token price (per 1M tokens). Higher means more capability per dollar. Models with no listed price are excluded.

Known limitations

Value rankings favor cheap models even if absolute performance is modest. A model scoring half as well at one-tenth the price wins on value — but may not meet your quality bar. Always check raw scores alongside value rankings.

About this ranking

Ranking data as of October 1, 2026

Agentic workloads are token-intensive — agents loop, retry, and chain multiple calls. That makes cost-per-token a critical factor alongside raw capability. This ranking divides each model's weighted agentic score (Terminal-Bench 2.0, BrowseComp, OSWorld-Verified) by its output token price. The result shows which models give you the most agent capability per dollar. If you're building production AI agents with budget constraints, this is where you start.

Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

MiMo-V2.6-Flash leads this ranking with a score of 195.29, followed by GPT-6 Luna (106.06) and MiMo-V2.6-Pro (76.37). There is a significant gap between the leading models and the rest of the field.

The best open-weight option is MiMo-V2.6-Flash (ranked #1 with a score of 195.29). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking uses provisional weighted averages across the scoring benchmarks in agentic. For detailed model profiles, click any model name above. To compare two specific models head-to-head, use the "vs #" links.

Explore More

Last updated: October 1, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.