Skip to main content
BenchLM
Data
Recommendation

Best LLMs for Writing in 2026

As of October 6, 2026, the top model in best llms for writing on the BenchLM leaderboard is MAI-Thinking-1 with a score of 94.7.

Bottom line: instruction following is what separates writing models in practice. Claude Fable 5 held the top human-preference rating in the September 2026 snapshot; the instruction-following leaders below execute a brief most faithfully.

Ranking data as of

Full Rankings (124 ranked)

1
MAI-Thinking-1
Microsoft·Proprietary·256K

94.7

prov. avg

2
Grok 4.3
xAI·Proprietary·1M

93.2

prov. avg

3
GPT-5.2-Codex
OpenAI·Proprietary·400K

92.4

prov. avg

4
GLM-5.1
Z.AI·Open Weight·203K

92.4

prov. avg

5
MiniMax M3
MiniMax·Open Weight·1M

92.4

prov. avg

6
MiMo-V2.5-Pro
Xiaomi·Proprietary·1M

92.4

prov. avg

7
GPT-5.5
OpenAI·Proprietary·1M

91.9

prov. avg

8
Muse Spark
Meta·Proprietary·262K

91.9

prov. avg

9
GPT-5.4 nano
OpenAI·Proprietary·400K

91.9

prov. avg

10
MiniMax M2.7
MiniMax·Open Weight·200K

91.6

prov. avg

11
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

91.6

prov. avg

12
Qwen3.5-27B
Alibaba·Open Weight·262K

91.5

prov. avg

13
Gemma 4 31B
Google·Open Weight·256K

91.5

prov. avg

14
GPT-5.3 Codex
OpenAI·Proprietary·400K

91.2

prov. avg

15
GPT-5.2
OpenAI·Proprietary·400K

91.2

prov. avg

16
Qwen3.8 Max
Alibaba·Open Weight·1M

89.8

prov. avg

17
Qwen3.7 Max
Alibaba·Proprietary·1M

89.2

prov. avg

18
Qwen3.7 Plus
Alibaba·Proprietary·1M

89.2

prov. avg

19
GPT-5.4
OpenAI·Proprietary·1.05M

89.2

prov. avg

20
Command A+
Cohere·Open Weight·128K

89.2

prov. avg

21
Gemma 4 12B
Google·Open Weight·256K

88.7

prov. avg

22
GLM-5.2
Z.AI·Open Weight·1M

88.5

prov. avg

23
Inkling-Small
Thinking Machines Lab·Open Weight·1M

88.5

prov. avg

24
GPT-5.4 mini
OpenAI·Proprietary·400K

88.5

prov. avg

25
GLM-5-Turbo
Z.AI·Proprietary·200K

88.3

prov. avg

26
GPT-5.1
OpenAI·Proprietary·200K

87.9

prov. avg

27
GPT-5.6 Sol
OpenAI·Proprietary·1.05M

87.7

prov. avg

28
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

87.4

prov. avg

29
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

87.4

prov. avg

30
Gemma 4 26B A4B
Google·Open Weight·256K

87.3

prov. avg

31
GLM-5
Z.AI·Open Weight·200K

87.2

prov. avg

32
Qwen3.8-Omni-Flash
Alibaba·Proprietary·1M

86.9

prov. avg

33
Qwen3.8-Flash-Next
Alibaba·Open Weight·262K

86.5

prov. avg

34
o3
OpenAI·Proprietary·200K

86

prov. avg

35
GPT-5.6 Terra
OpenAI·Proprietary·1.05M

85.7

prov. avg

36
Solar Pro 3
Upstage·Proprietary·128K

85.7

prov. avg

37
GPT-5 (medium)
OpenAI·Proprietary·128K

84.9

prov. avg

38
Gemini 3 Pro
Google·Proprietary·2M

84.7

prov. avg

39
dots3-note Preview
Dots Studio·Open Weight·512K

84.5

prov. avg

40
o1
OpenAI·Proprietary·200K

84.5

prov. avg

41
Kimi K2.5
Moonshot AI·Open Weight·256K

84.4

prov. avg

42
GPT-5.1-Codex
OpenAI·Proprietary·400K

84.2

prov. avg

43
Gemini 3.5 Flash
Google·Proprietary·1M

84.1

prov. avg

44
Inkling
Thinking Machines Lab·Open Weight·1M

83.2

prov. avg

45
GPT-OSS 120B
OpenAI·Open Weight·128K

82.8

prov. avg

46
MiMo-V2-Pro
Xiaomi·Proprietary·1M

82.6

prov. avg

47
Mistral Medium 3.5 128B
Mistral·Open Weight·256K

82.6

prov. avg

48
Qwen3.6 Plus
Alibaba·Proprietary·1M

82.5

prov. avg

49
Qwen3.8-27B
Alibaba·Open Weight·262K

82.5

prov. avg

50
Granite 4.2 8B
IBM·Open Weight·128K

82.1

prov. avg

51
GLM-4.7
Z.AI·Open Weight·200K

81.4

prov. avg

52
Qwen3.6-27B
Alibaba·Open Weight·262K

81

prov. avg

53
Step 3.7 Flash
StepFun·Open Weight·256K

80.6

prov. avg

54
GPT-OSS 20B
OpenAI·Open Weight·128K

77.8

prov. avg

55
Muse Glimmer 30B
Meta·Open Weight·131K

77

prov. avg

56
Mercury 2.5
Inception·Proprietary·260K

77

prov. avg

57
Qwen3.6-35B-A3B
Alibaba·Open Weight·262K

76.8

prov. avg

58
Claude Fable 5
Anthropic·Proprietary·1M+

75.7

prov. avg

59
Kimi K2.7 Code
Moonshot AI·Open Weight·256K

75.1

prov. avg

60
Claude Opus 4.8
Anthropic·Proprietary·1M

74

prov. avg

61
GLM-5V-Turbo
Z.AI·Proprietary·200K

72.5

prov. avg

62
Nemotron 3 Nano Omni 30B A3B
NVIDIA·Open Weight·256K

72.4

prov. avg

63
Ling 3.0 Flash
InclusionAI·Open Weight·262K

71.4

prov. avg

64
Ternary Bonsai 2 27B
Prism ML·Open Weight·262K

70.3

prov. avg

65
Ling 3.0 Flash FP8
InclusionAI·Open Weight·262K

69

prov. avg

66
Claude Opus 4.5 Thinking
Anthropic·Proprietary·200K

68.5

prov. avg

67

67.8

prov. avg

68
Trinity-Large-Thinking
Arcee AI·Open Weight·512K

66.3

prov. avg

69
Claude 4.1 Opus Thinking
Anthropic·Proprietary·200K

65.1

prov. avg

70
Gemini 3 Flash
Google·Proprietary·1M

64.7

prov. avg

71
Grok 4
xAI·Proprietary·128K

62.9

prov. avg

72
MiMo-V2-Omni
Xiaomi·Proprietary·262K

62.6

prov. avg

73
Grok 4.1 Fast (Reasoning)
xAI·Proprietary·2M

61.6

prov. avg

74
Grok 4 Fast (Reasoning)
xAI·Proprietary·2M

58.7

prov. avg

75
DeepSeek V3.2
DeepSeek·Open Weight·128K

56.7

prov. avg

76
Gemini 2.5 Pro
Google·Proprietary·1M

56.4

prov. avg

77
Mistral Small 4
Mistral·Open Weight·256K

55.7

prov. avg

78
Claude 4 Sonnet
Anthropic·Proprietary·200K

52

prov. avg

79
Claude Opus 4.6
Anthropic·Proprietary·1M

51

prov. avg

80
Gemma 4 E4B
Google·Open Weight·128K

50.5

prov. avg

81
Qwen3 Max
Alibaba·Proprietary·1M

50.3

prov. avg

82
GPT-4.1
OpenAI·Proprietary·1M

48.9

prov. avg

83
Llama 4 Maverick
Meta·Open Weight·1M

48.9

prov. avg

84
DeepSeek V3.1 (Reasoning)
DeepSeek·Open Weight·128K

47

prov. avg

85
Kimi K2
Moonshot AI·Proprietary·128K

47

prov. avg

86
Grok Code Fast 1
xAI·Proprietary·256K

46.8

prov. avg

87
Claude Sonnet 4.6
Anthropic·Proprietary·200K

46.6

prov. avg

88
DeepSeek V3 0324
DeepSeek·Open Weight·128K

46.3

prov. avg

89
Hy3 Preview
Tencent·Open Weight·256K

46.1

prov. avg

90
MiMo-V2-Flash
Xiaomi·Open Weight·256K

44.9

prov. avg

91
DeepSeek-R1
DeepSeek·Open Weight·128K

44.5

prov. avg

92
Llama 4 Scout
Meta·Open Weight·10M

44.3

prov. avg

93
Mistral Medium 3
Mistral·Proprietary·128K

44.1

prov. avg

94
Gemini 2.5 Flash
Google·Proprietary·1M

43.7

prov. avg

95
GPT-4.1 mini
OpenAI·Proprietary·1M

42.8

prov. avg

96
Nova Pro
Amazon·Proprietary·128K

42.5

prov. avg

97
Gemma 4 E2B
Google·Open Weight·128K

42.4

prov. avg

98
DeepSeek V3.1
DeepSeek·Open Weight·128K

42.1

prov. avg

99
GLM-4.5-Air
Z.AI·Proprietary·128K

41.9

prov. avg

100
GLM-4.6
Z.AI·Open Weight·200K

40.7

prov. avg

101
Grok 4.1 Fast
xAI·Proprietary·1M

40.4

prov. avg

102
Mistral Large 3
Mistral·Proprietary·128K

40

prov. avg

103
Claude 3 Haiku
Anthropic·Proprietary·200K

39.9

prov. avg

104
DeepSeek V3
DeepSeek·Open Weight·128K

38.2

prov. avg

105
Sarvam 105B
Sarvam·Open Weight·128K

37.7

prov. avg

106
GPT-4o
OpenAI·Proprietary·128K

37.6

prov. avg

107
Ling 2.6 Flash
InclusionAI·Open Weight·262K

37.5

prov. avg

108
LFM2.5-2.6B
LiquidAI·Open Weight·128K

37.4

prov. avg

109
LFM2.5-8B-A1B
LiquidAI·Open Weight·128K

35.4

prov. avg

110
GPT-4.1 nano
OpenAI·Proprietary·1M

34.6

prov. avg

111
Gemma 3 27B
Google·Open Weight·32K

34.3

prov. avg

112
Mistral Large 2
Mistral·Proprietary·128K

33.5

prov. avg

113
Qwen3-Omni-30B-A3B-Instruct
Alibaba·Open Weight·N/A

33.5

prov. avg

114
GPT-4o mini
OpenAI·Proprietary·128K

33.2

prov. avg

115
Claude Opus 4.5
Anthropic·Proprietary·200K

30.6

prov. avg

116
DeepSeek R1 Distill Qwen 32B
DeepSeek·Open Weight·128K

27.4

prov. avg

117
Granite-4.0-H-1B
IBM·Open Weight·128K

27.4

prov. avg

118
Sarvam 30B
Sarvam·Open Weight·64K

27.4

prov. avg

119
Granite-4.0-350M
IBM·Open Weight·32K

27.4

prov. avg

120
Granite-4.0-H-350M
IBM·Open Weight·32K

27.4

prov. avg

121
Phi-4
Microsoft·Open Weight·16K

27.4

prov. avg

122
ZAYA1-8B
Zyphra·Open Weight·131K

22.8

prov. avg

123
MiniCPM5-1B
OpenBMB·Open Weight·131K

9.7

prov. avg

124
LFM2.5-230M
LiquidAI·Open Weight·32K

9.7

prov. avg

Current position

MAI-Thinking-1 leads the live llms for writing ranking at 94.7.

Grok 4.3 ranks #2 at 93.2.

GPT-5.2-Codex ranks #3 at 92.4.

How to choose

Key Takeaways

The top model is MAI-Thinking-1 by Microsoft with a provisional score of 94.7.

The best open-weight model is GLM-5.1 at position #4.

124 models are included in this ranking.

Score in Context

What these scores mean

Writing quality has no direct benchmark, so this page ranks by instruction following — whether the model does what the brief asked. Read it with a blind human-preference rating from an independent index for prose quality.

Known limitations

Style is subjective and prompt-sensitive. Instruction-following scores reward constraint compliance, not voice — a model can follow your brief perfectly and still write flat prose. Test your actual editing workflow.

About this ranking

Ranking data as of October 6, 2026

There is no single "writing benchmark," so BenchLM ranks writing capability by the instruction-following category — the best available proxy for whether a model matches your tone, structure, and length constraints — read alongside human-preference ratings from an independent index. Claude models held the top of that index in the September 2026 snapshot; the instruction-following table shows who follows a brief most reliably.

Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

MAI-Thinking-1 leads this ranking with a score of 94.7, followed by Grok 4.3 (93.2) and GPT-5.2-Codex (92.4). The top three are separated by just a few points — any of them would perform well for this use case.

The best open-weight option is GLM-5.1 (ranked #4 with a score of 92.4). While proprietary models lead, open-weight options are within striking distance for teams willing to trade a few points of performance for full model control.

This ranking uses provisional weighted averages across the scoring benchmarks in instructionFollowing. For detailed model profiles, click any model name above. To compare two specific models head-to-head, use the "vs #" links.

Questions

What is the best LLM for writing?

For prose, our pick is Claude Fable 5: it held the highest human-preference rating we track in the September 2026 snapshot, the best available signal for how writing reads. For structured copy with strict format and length rules, start instead from the leaders of the instruction-following table above, which recomputes on every build.

What is the best LLM for creative writing?

Human preference is the best available signal for creative prose, and Claude models led the September 2026 snapshot of the independent index we track, with Claude Fable 5 on top. Non-reasoning models often feel more natural for iterative drafting because they respond without a thinking pause.

What is the best AI for resume writing?

Resume writing is an instruction-following task: strict format, tight length, specific tone. Any model near the top of the table above handles it well, so the practical differences are price and speed. Pick the cheapest row that clears about 90 on the price-performance page, then test it on your own template.

Are benchmarks meaningful for writing quality?

Partially. Instruction following measures whether the model obeyed the brief — essential for professional writing — and an independent human-preference rating captures blind reader preference. Neither measures your voice. Use the scores to shortlist, then run a 10-prompt bake-off in your own editing workflow.

Explore More

Last updated: October 6, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.