Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start the free Radar Brief
BenchLM recommendation

Best LLMs for Roleplay in 2026

Data verified

As of August 30, 2026, the top model in best llms for roleplay on the BenchLM leaderboard is Qwen3.5-27B with a score of 95.

Bottom line: persona consistency is instruction following in costume — the IFEval/IFBench leaders below hold character best. For prose feel, cross-check Arena Elo and test your own scenarios.

About this ranking

Last verified: August 30, 2026

There is no dedicated roleplay benchmark, so this reporting family ranks the measurable ingredients: instruction following (IFEval, IFBench — whether the model stays in the persona and format you set), WildBench (performance on real, messy user prompts, including creative ones), and MuSR (tracking multi-step narrative state). Human preference for prose style is worth reading alongside: Claude models hold the top Arena Elo scores BenchLM tracks.

This page ranks models by a sourced proxy blend — instruction following, real-prompt quality, and narrative reasoning — because no standalone roleplay benchmark exists.

Qwen3.5-27B leads this ranking with a score of 95, followed by Agents-A1 (94.8) and Kimi K2.5 (93.9). The top three are separated by just a few points — any of them would perform well for this use case.

The best open-weight option is Qwen3.5-27B (ranked #1 with a score of 95). Open-weight models are highly competitive in this category — self-hosting is a viable alternative to proprietary APIs.

This ranking uses provisional overall weighted scores from the active scoring formula. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

How to choose

Full Rankings (43 models)

1
Qwen3.5-27B
Alibaba·Open Weight·262K

95

sourced avg

2
Agents-A1
InternScience·Open Weight·262K

94.8

sourced avg

3
Kimi K2.5
Moonshot AI·Open Weight·256K

93.9

sourced avg

4
o3-mini
OpenAI·Proprietary·200K

93.9

sourced avg

5
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

93.4

sourced avg

6
GLM-5
Z.AI·Open Weight·200K

92.6

sourced avg

7
Qwen3.5 397B
Alibaba·Open Weight·128K

92.6

sourced avg

8
o1
OpenAI·Proprietary·200K

92.2

sourced avg

9
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

91.9

sourced avg

10
GPT-4.1 mini
OpenAI·Proprietary·1M

88.5

sourced avg

11
GPT-4.1
OpenAI·Proprietary·1M

87.4

sourced avg

12
dots3-note Preview
Dots Studio·Open Weight·512K

87.2

sourced avg

13
Qwen3.7 Plus
Alibaba·Proprietary·1M

86.8

sourced avg

14
Qwen3.7 Max
Alibaba·Proprietary·1M

86.7

sourced avg

15
DeepSeek V3
DeepSeek·Open Weight·128K

86.1

sourced avg

16
Qwen3.6 Plus
Alibaba·Proprietary·1M

85.1

sourced avg

17
MAI-Thinking-1
Microsoft·Proprietary·256K

85

sourced avg

18
GPT-4.1 nano
OpenAI·Proprietary·1M

83.2

sourced avg

19
Qwen3.8 Max
Alibaba·Open Weight·1M

82.8

sourced avg

20
Inkling-Small
Thinking Machines Lab·Open Weight·1M

82.2

sourced avg

21
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

81.7

sourced avg

22
Grok 4.3
xAI·Proprietary·1M

81.3

sourced avg

23
Qwen3.8-Flash-Next
Alibaba·Open Weight·262K

81.3

sourced avg

24
Celeris-1
Celeris·Proprietary·128K

80.8

sourced avg

25
Inkling
Thinking Machines Lab·Open Weight·1M

79.8

sourced avg

26
Qwen3.8-27B
Alibaba·Open Weight·262K

79.5

sourced avg

27
Muse Glimmer 30B
Meta·Open Weight·131K

77

sourced avg

28
Mellum2-12B-A2.5B-Thinking
JetBrains·Open Weight·128K

76.5

sourced avg

29
Gemini 3.5 Flash
Google·Proprietary·1M

76.3

sourced avg

30
Mellum2-12B-A2.5B-Instruct
JetBrains·Open Weight·128K

75.8

sourced avg

31
Claude Opus 4.5
Anthropic·Proprietary·200K

74.5

sourced avg

32
Ling 3.0 Flash
InclusionAI·Open Weight·262K

74.5

sourced avg

33
LFM2.5-8B-A1B
LiquidAI·Open Weight·128K

74.2

sourced avg

34
Nemotron 3 Nano Omni 30B A3B
NVIDIA·Open Weight·256K

74.2

sourced avg

35
Ling 3.0 Flash FP8
InclusionAI·Open Weight·262K

73.4

sourced avg

36

72.9

sourced avg

37
ZAYA1-8B
Zyphra·Open Weight·131K

69.1

sourced avg

38
MiniCPM5-1B
OpenBMB·Open Weight·131K

63.5

sourced avg

39
Hy3 Preview
Tencent·Open Weight·256K

63.1

sourced avg

40
LFM2.5-VL-450M
LiquidAI·Open Weight·128K

61.2

sourced avg

41
LFM2.5-2.6B
LiquidAI·Open Weight·128K

59.2

sourced avg

42
Ling 2.6 Flash
InclusionAI·Open Weight·262K

57

sourced avg

43
LFM2.5-230M
LiquidAI·Open Weight·32K

55.1

sourced avg

Key Takeaways

The top model on this sourced reporting-family slice is Qwen3.5-27B by Alibaba with an average of 95.

The best open-weight model is Qwen3.5-27B at position #1.

43 models are listed with sourced benchmark coverage in this reporting family.

Score in Context

What these scores mean

A proxy blend: instruction following measures whether the model keeps your persona, format, and constraints; WildBench samples real user prompts; MuSR tests narrative-state tracking. Together they approximate roleplay reliability.

Known limitations

No benchmark measures prose charm or character voice — the qualities roleplay users often care about most. Content policies also differ sharply between providers and matter more than a few benchmark points for this use case. Test your actual scenarios.

Best LLMs for Roleplay FAQ

What is the best LLM for roleplay?

By measurable proxies — instruction adherence, real-prompt quality, narrative tracking — the top rows of this table lead. For prose feel, Claude models hold the highest human-preference (Arena Elo) scores BenchLM tracks. The honest answer is to shortlist from this table and run your own scenarios; voice preference is personal.

What is the best local LLM for roleplay?

The strongest open-weight rows in this family can run locally — self-hosting also gives you full control over system prompts and content settings, which matters for this use case. See the local LLM rankings for what fits your VRAM tier, and the Ollama guide for pull commands.

Why is there no roleplay benchmark?

Because the target is subjective: persona charm and voice resist automated scoring. What can be measured — does the model follow persona instructions, handle real messy prompts, and track narrative state — is what this family blends. Treat it as a reliability floor, not a style ranking.

Do content policies matter more than benchmarks for roleplay?

Often yes. Providers differ sharply in what fiction they permit and how aggressively safety layers interrupt scenes, and that difference dwarfs a few benchmark points in practice. Check each provider's current usage policies for your use case; open-weight models under your own hosting sidestep the issue entirely.

Last updated: August 30, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.