Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

BenchLM recommendation

Best LLMs for Writing in 2026

Data verified

As of August 20, 2026, the top model in best llms for writing on the BenchLM leaderboard is Qwen3.8 Max with a score of 97.9.

Last verified: August 20, 2026

There is no single "writing benchmark," so BenchLM ranks writing capability by the instruction-following category — the best available proxy for whether a model matches your tone, structure, and length constraints — read alongside Arena Elo, the human-preference signal. Claude models hold the top Arena Elo scores; the instruction-following table below shows who follows a brief most reliably.

Unless noted otherwise, ranking surfaces on this page use BenchLM's provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

Bottom line: instruction following is what separates writing models in practice. Claude Fable 5 holds the top Arena Elo (1508) for human preference; the instruction-following leaders below execute a brief most faithfully.

Qwen3.8 Max leads this ranking with a score of 97.9, followed by MAI-Thinking-1 (97.9) and Inkling-Small (96.5). The top three are separated by just a few points — any of them would perform well for this use case.

The best open-weight option is Qwen3.8 Max (ranked #1 with a score of 97.9). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking uses provisional weighted averages across the scoring benchmarks in instructionFollowing. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

What changed

MAI-Thinking-1 leads the instruction-following category at 93.3.

GPT-5.4 close second at 92.6 — the reliable structured-writing workhorse.

Claude Opus 4.6 strong instruction following (91.3) with Claude's prose quality.

How to choose

Full Rankings (40 models)

1
Qwen3.8 Max
Alibaba·Open Weight·1M

97.9

prov. avg

2
MAI-Thinking-1
Microsoft·Proprietary·256K

97.9

prov. avg

3
Inkling-Small
Thinking Machines Lab·Open Weight·1M

96.5

prov. avg

4
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

95.4

prov. avg

5
Grok 4.3
xAI·Proprietary·1M

94.5

prov. avg

6
dots3-note Preview
Dots Studio·Open Weight·512K

94.3

prov. avg

7
Qwen3.5-27B
Alibaba·Open Weight·262K

93.7

prov. avg

8
Agents-A1
InternScience·Open Weight·262K

93.7

prov. avg

9
Qwen3.7 Plus
Alibaba·Proprietary·1M

93

prov. avg

10
Qwen3.7 Max
Alibaba·Proprietary·1M

92.6

prov. avg

11
Inkling
Thinking Machines Lab·Open Weight·1M

91.1

prov. avg

12
Kimi K2.5
Moonshot AI·Open Weight·256K

91

prov. avg

13
o3-mini
OpenAI·Proprietary·200K

91

prov. avg

14
Qwen3.8-27B
Alibaba·Open Weight·262K

90.4

prov. avg

15
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

89.6

prov. avg

16
GLM-5
Z.AI·Open Weight·200K

87.2

prov. avg

17
Qwen3.5 397B
Alibaba·Open Weight·128K

87.2

prov. avg

18
Qwen3.6 Plus
Alibaba·Proprietary·1M

86.6

prov. avg

19
o1
OpenAI·Proprietary·200K

86.1

prov. avg

20
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

85.2

prov. avg

21
Muse Glimmer 30B
Meta·Open Weight·131K

84.7

prov. avg

22
Gemini 3.5 Flash
Google·Proprietary·1M

83.1

prov. avg

23
Ling 3.0 Flash
InclusionAI·Open Weight·262K

79

prov. avg

24
Nemotron 3 Nano Omni 30B A3B
NVIDIA·Open Weight·256K

78.3

prov. avg

25
Ling 3.0 Flash FP8
InclusionAI·Open Weight·262K

76.4

prov. avg

26

75.3

prov. avg

27
GPT-4.1 mini
OpenAI·Proprietary·1M

75.2

prov. avg

28
GPT-4.1
OpenAI·Proprietary·1M

72

prov. avg

29
GPT-4.1 nano
OpenAI·Proprietary·1M

59.8

prov. avg

30
Hy3 Preview
Tencent·Open Weight·256K

52.9

prov. avg

31
Claude Opus 4.5
Anthropic·Proprietary·200K

49.4

prov. avg

32
LFM2.5-8B-A1B
LiquidAI·Open Weight·128K

48.1

prov. avg

33
LFM2.5-2.6B
LiquidAI·Open Weight·128K

43.9

prov. avg

34
Mellum2-12B-A2.5B-Thinking
JetBrains·Open Weight·128K

40.2

prov. avg

35
Ling 2.6 Flash
InclusionAI·Open Weight·262K

39

prov. avg

36
Mellum2-12B-A2.5B-Instruct
JetBrains·Open Weight·128K

38.1

prov. avg

37
ZAYA1-8B
Zyphra·Open Weight·131K

31.6

prov. avg

38
LFM2.5-VL-450M
LiquidAI·Open Weight·128K

26.2

prov. avg

39
MiniCPM5-1B
OpenBMB·Open Weight·131K

13.2

prov. avg

40
LFM2.5-230M
LiquidAI·Open Weight·32K

1

prov. avg

Key Takeaways

The top model is Qwen3.8 Max by Alibaba with a provisional score of 97.9.

The best open-weight model is Qwen3.8 Max at position #1.

40 models are included in this ranking.

Score in Context

What these scores mean

Writing quality has no direct benchmark, so this page ranks by instruction following — whether the model does what the brief asked. Read it with Arena Elo, the blind human-preference score, for prose quality.

Known limitations

Style is subjective and prompt-sensitive. Instruction-following scores reward constraint compliance, not voice — a model can follow your brief perfectly and still write flat prose. Test your actual editing workflow.

Best LLMs for Writing FAQ

What is the best LLM for writing?

Claude Fable 5 is the strongest overall pick — it holds the highest Arena Elo BenchLM tracks (1508), the best human-preference signal for prose, with top-tier instruction following. For structured copy at lower cost, GPT-5.4's 92.6 instruction-following score makes it the workhorse choice.

What is the best LLM for creative writing?

Human preference is the best available signal for creative prose, and Claude models lead it: Claude Fable 5 holds the top Arena Elo on BenchLM's board. Non-reasoning models often feel more natural for iterative drafting because they respond without a thinking pause.

What is the best AI for resume writing?

Resume writing is an instruction-following task — strict format, tight length, specific tone. Any model above ~90 in the table handles it well; the practical differences are price and speed, so a mid-tier model like GPT-5.4 or Claude Sonnet 5 is the sensible default.

Are benchmarks meaningful for writing quality?

Partially. Instruction following measures whether the model obeyed the brief — essential for professional writing — and Arena Elo captures blind human preference. Neither measures your voice. Use the scores to shortlist, then run a 10-prompt bake-off in your own editing workflow.

Last updated: August 20, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.