Skip to main content

Model profile

Qwen3.5 397B

AlibabaCurrentReleased Feb 16, 2026
Data verified
Overall Score
57.01Public #71 of 200
Arena Elo
1400
Eligible category ranks
6of 8
Price (1M tokens)
$0.6 in / $3.6 out
API pricing
Speed
96tok/s
Context
128K

Evidence coverage

55 of 321 tracked benchmarks are published. 38 are verified and 17 provisional. 8 of 8 categories are measured.

Updated July 20, 2026Methodology
Published / tracked
55 / 321
Verified
38
Provisional
17
Categories with evidence
8 / 8

Evidence by category

  • Agentic18 benchmarks
    Mixed evidence
  • Coding5 benchmarks
    Mixed evidence
  • Reasoning4 benchmarks
    Mixed evidence
  • Knowledge12 benchmarks
    Mixed evidence
  • Math5 benchmarks
    Verified
  • Multilingual2 benchmarks
    Verified
  • Multimodal7 benchmarks
    Mixed evidence
  • Inst. Following2 benchmarks
    Mixed evidence
Open WeightSelf-hostNon-Reasoning
Confidence:
Very high
base

Qwen3.5 397B ranks #71 out of 200 models on the public leaderboard with an overall score of 57.01/100. It does not yet have enough sourced coverage for BenchLM's verified leaderboard. While not a frontier model, it offers specific advantages depending on the use case.

Qwen3.5 397B is a open weight model with a 128K token context window. It processes queries without explicit chain-of-thought reasoning, offering faster response times and lower token usage.

Qwen3.5 397B sits inside the Qwen3.5 397B family alongside Qwen3.5 397B (Reasoning). This profile currently has 55 of 321 tracked benchmarks. BenchLM only exposes non-generated benchmark rows publicly, so missing categories stay blank until a sourced evaluation is available.

Its strongest category is Multilingual (#5), while its weakest is Agentic (#67). This performance profile makes it a well-rounded choice across a range of tasks.

Peer position

Exact provisional scores and ranks for the closest listed peers. A score can appear before a model clears the evidence threshold for a rank, so equal scores can have different rank states.

Range 56.5957.25

  1. Gemini 2.5 Pro
    Google
    #7057.25
    Gemini 2.5 Pro is #70 with a score of 57.25.
    Compare
  2. Qwen3.5 397BCurrent model
    Alibaba
    #7157.01
    Qwen3.5 397B is #71 with a score of 57.01.
  3. Qwen3.5-35B-A3B
    Alibaba
    #7256.97
    Qwen3.5-35B-A3B is #72 with a score of 56.97.
    Compare
  4. GPT-5.3-Codex-Spark
    OpenAI
    #7356.91
    GPT-5.3-Codex-Spark is #73 with a score of 56.91.
    Compare
  5. Kimi K2.6
    Moonshot AI
    #7456.79
    Kimi K2.6 is #74 with a score of 56.79.
    Compare
  6. GPT-5.4 mini
    OpenAI
    #7556.77
    GPT-5.4 mini is #75 with a score of 56.77.
    Compare
  7. Grok 4 Fast (Reasoning)
    xAI
    #7656.59
    Grok 4 Fast (Reasoning) is #76 with a score of 56.59.
    Compare

Category percentile

More

Relative position among models eligible for each sourced category. A higher percentile means a stronger position within that category's ranked cohort; 100 is highest.

  1. Multilingual67%
    Eligible cohort rank #5 of 13Category score 69.7
  2. Inst. Following65%
    Eligible cohort rank #14 of 38Category score 86.5
  3. Multimodal43%
    Eligible cohort rank #17 of 29Category score 66.4
  4. Knowledge20%
    Eligible cohort rank #42 of 52Category score 57.5
  5. Coding62%
    Eligible cohort rank #47 of 122Category score 52.9
  6. Agentic44%
    Eligible cohort rank #67 of 119Category score 46.1

Category evidence

Scores and ranks appear only where this model has published benchmark evidence. Categories without displayable source records remain not measured.

Category scores, ranks, weighting, benchmark coverage, and evidence status
CategoryScore
AgenticRank #67 of 119Percentile 44thWeight 22%18 benchmarksMixed sources46.1
CodingRank #47 of 122Percentile 62ndWeight 20%5 benchmarksMixed sources52.9
ReasoningRank Not rankedWeight 17%4 benchmarksMixed sources87.3
KnowledgeRank #42 of 52Percentile 20thWeight 12%12 benchmarksMixed sources57.5
MathRank Not rankedWeight 5%5 benchmarksVerified74.7
MultilingualRank #5 of 13Percentile 67thWeight 7%2 benchmarksVerified69.7
MultimodalRank #17 of 29Percentile 43rdWeight 12%7 benchmarksMixed sources66.4
Inst. FollowingRank #14 of 38Percentile 65thWeight 5%2 benchmarksMixed sources86.5

Chatbot Arena performance

Scroll horizontally to inspect confidence intervals and vote counts.

Chatbot Arena Elo, confidence interval, and vote count by evaluation view
ViewEloConfidence intervalVotes
Text Overall1400Not availableNot available

Benchmark Details

Rows below have a displayable published verification record. Each source link and provenance note remains in the page HTML while its category is closed. Source-unverified manual rows and generated rows stay hidden.

Agentic18 benchmarks
Terminal-Bench 2.0Provider exact
52.5%Weighted 38%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
BrowseCompProvider exact
62%Weighted 28%
Source: Qwen3.5-397B-A17B model cardProvenance: Provider exact
Claw-EvalBenchmark exact
56.8%Display only
Source: Claw-Eval leaderboardProvenance: Claw-Eval reports this model as qwen3.5-397b-a17b in the official 2026-05-09 leaderboard snapshot. BenchLM stores the primary Pass^3 value on the local Claw-Eval display key.
QwenClawBenchProvider exact
51.8%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
τ³-bench resultsProvider exact

τ³-Bench Tool-Agent-User Evaluation

68.4%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
VITA-BenchProvider exact
43.7%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
DeepPlanningProvider exact
37.6%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
ToolathlonProvider exact
36.3%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
MCP AtlasProvider exact
46.1%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
MCP-TasksProvider exact
74.2%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
WideResearchProvider exact
74.0%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
τ²-bench resultsReported

τ²-Bench Tool-Agent-User Evaluation

95.6%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Gert LabsBenchmark exact

Gert Labs Composite Game Benchmark

46.76%Display only
Source: Gert Labs rankingsProvenance: Gert Labs reports this composite leaderboard score in the public rankings API. BenchLM scales the source gscore from 0-1 to 0-100 and stores it as a display-only agentic benchmark.
ResearchClawBenchBenchmark exact
14.2%Display only
Source: ResearchClawBench leaderboardProvenance: ResearchClawBench reports this model as ResearchHarness (Qwen3.5-397B-A17B) in the official Pass@1 leaderboard. BenchLM stores the one-decimal RADS average on the local ResearchClawBench display key and excludes it from weighted rankings.
AA Agentic IndexReported

Artificial Analysis Agentic Index

19.9%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
APEX-Agents-AAReported
15.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
GDPval-AAReported

GDPval-AA normalized

23.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
GDPval-AAReported
962Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Coding5 benchmarks
SWE-bench VerifiedProvider exact

Software Engineering Benchmark Verified

76.2%Weighted 16%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
SWE-bench ProProvider exact
50.9%Weighted 10%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
LiveCodeBench v6Provider exact
83.6%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
AA-SciCodeReported

Artificial Analysis SciCode

42.0%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA Coding IndexReported

Artificial Analysis Coding Index

48.2%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Reasoning4 benchmarks
LongBench v2Provider exact
63.2%Weighted 38%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
AI-NeedleProvider exact
68.7%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
AA-LCRReported

Artificial Analysis Long Context Reasoning

65.7%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
CritPtReported

Critical Physics Tasks

1.7%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Knowledge12 benchmarks
HLEProvider exact

Humanity's Last Exam

28.7%Weighted 45%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
MMLU-ProProvider exact

Massive Multitask Language Understanding Professional

87.8%Weighted 30%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
GPQAProvider exact

Graduate-Level Google-Proof Q&A

88.4%Weighted 7%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
SuperGPQAProvider exact

SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines

70.4%Weighted 7%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
MMLU-ReduxProvider exact
94.9%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
C-EvalProvider exact
93%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
Artificial Analysis Intelligence IndexReported
33.7%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-GPQA DiamondReported

Artificial Analysis GPQA Diamond

89.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-HLEReported

Artificial Analysis Humanity's Last Exam

27.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience IndexReported

Artificial Analysis Omniscience Index

-29.8%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience AccuracyReported

Artificial Analysis Omniscience Accuracy

31.4%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
AA-Omniscience Hallucination RateReported

Artificial Analysis Omniscience Hallucination Rate

89.1%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Math5 benchmarks
AIME26Provider exact

AIME 2026

93.3%Weighted 25%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
HMMT Feb 2026Provider exact

Harvard-MIT Mathematics Tournament February 2026

87.9%Weighted 25%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
HMMT Feb 2025Provider exact

Harvard-MIT Mathematics Tournament February 2025

94.8%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
HMMT Nov 2025Provider exact

Harvard-MIT Mathematics Tournament November 2025

92.7%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
MMAnswerBenchProvider exact
80.9%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
Multilingual2 benchmarks
MMLU-ProXProvider exact
84.7%Weighted 100%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
NOVA-63Provider exact
59.1%Display only
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
Multimodal7 benchmarks
MMMU-ProProvider exact

Massive Multi-discipline Multimodal Understanding Pro

79%Weighted 45%
Source: Qwen: Qwen3.6-Plus multimodal comparison tableProvenance: Provider exact
CharXivProvider exact

CharXiv Reasoning

80.8%Weighted 25%
Source: Qwen: Qwen3.6-Plus multimodal comparison tableProvenance: Provider exact
MathVisionProvider exact
88.6%Display only
Source: Qwen: Qwen3.6-Plus multimodal comparison tableProvenance: Provider exact
VideoMMMUProvider exact
84.7%Display only
Source: Qwen: Qwen3.6-Plus multimodal comparison tableProvenance: Provider exact
ScreenSpot ProProvider exact
65.6%Display only
Source: Qwen: Qwen3.6-Plus multimodal comparison tableProvenance: Provider exact
V*Provider exact
95.8%Display only
Source: Qwen: Qwen3.6-Plus multimodal comparison tableProvenance: Qwen reports V* as with-CI / without-CI. BenchLM stores the first published value.
AA-MMMU-ProReported

Artificial Analysis MMMU-Pro

77.3%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.
Inst. Following2 benchmarks
IFEvalProvider exact

Instruction-Following Eval

92.6%Weighted 35%
Source: Qwen: Qwen3.6-Plus comparison tableProvenance: Provider exact
AA-IFBenchReported

Artificial Analysis IFBench

78.8%Display only
Source: Artificial Analysis model benchmarksProvenance: Display-only row synced from the current Artificial Analysis model payload. It is excluded from BenchLM weighted scoring.

Qwen3.5 397B Family

Base entry

Frequently Asked Questions

How does Qwen3.5 397B perform overall in AI benchmarks?

Qwen3.5 397B currently ranks #71 out of 200 models on BenchLM's provisional leaderboard with an overall score of 57.01. It is created by Alibaba. Its published context window is 128K.

Is Qwen3.5 397B good for knowledge and understanding?

Qwen3.5 397B ranks #42 out of 52 models in knowledge and understanding benchmarks with an average score of 57.5. There are stronger options in this category.

Is Qwen3.5 397B good for coding and programming?

Qwen3.5 397B ranks #47 out of 122 models in coding and programming benchmarks with an average score of 52.9. There are stronger options in this category.

Is Qwen3.5 397B good for mathematics?

Qwen3.5 397B has visible benchmark coverage in mathematics, but BenchLM does not currently assign it a global category rank there.

Is Qwen3.5 397B good for reasoning and logic?

Qwen3.5 397B has visible benchmark coverage in reasoning and logic, but BenchLM does not currently assign it a global category rank there.

Is Qwen3.5 397B good for agentic tool use and computer tasks?

Qwen3.5 397B ranks #67 out of 119 models in agentic tool use and computer tasks benchmarks with an average score of 46.1. There are stronger options in this category.

Is Qwen3.5 397B good for multimodal and grounded tasks?

Qwen3.5 397B ranks #17 out of 29 models in multimodal and grounded tasks benchmarks with an average score of 66.4. There are stronger options in this category.

Is Qwen3.5 397B good for instruction following?

Qwen3.5 397B ranks #14 out of 38 models in instruction following benchmarks with an average score of 86.5. There are stronger options in this category.

Is Qwen3.5 397B good for multilingual tasks?

Qwen3.5 397B ranks #5 out of 13 models in multilingual tasks benchmarks with an average score of 69.7. It is among the top performers in this category.

Is Qwen3.5 397B open source?

Yes, Qwen3.5 397B is an open weight model created by Alibaba, meaning it can be downloaded and run locally or fine-tuned for specific use cases.

Which sibling models are related to Qwen3.5 397B?

Qwen3.5 397B belongs to the Qwen3.5 397B family. Related variants on BenchLM include Qwen3.5 397B (Reasoning).

Does Qwen3.5 397B have full benchmark coverage on BenchLM?

Not yet. Qwen3.5 397B currently has 55 published benchmark scores out of the 321 benchmarks BenchLM tracks. BenchLM only exposes non-generated public benchmark rows, so missing categories stay blank until a sourced evaluation is available.

What is the context window size of Qwen3.5 397B?

Qwen3.5 397B has a published context window of 128K, which determines how much text it can process in a single interaction.

Last updated: July 20, 2026 · Runtime metrics stay blank until BenchLM has a sourced snapshot.

Choose with this week’s evidence

Join 2,000+ readers for ranking moves, new releases, pricing changes, and the evidence behind them.

Free. One email per week.