Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Vals SkillsBench (SkillsBench)

We mirror this table; we do not rank on it.

Data verified 35 confirmed releases in the last 30 daysFollow model changes

How important are skills for agents?

SkillsBench score on SkillsBench — September 11, 2026

We mirror the published skillsbench score view for SkillsBench. DeepSeek V4.1 Flash leads the public snapshot at 69.80%, followed by Grok 4.5 (66.03%) and Gemini 3.7 Flash (65.89%). We do not use these results to rank models overall.

35 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated September 11, 2026

SkillsBench score table (35 models)

Score
1
DeepSeek V4.1 FlashDeepSeek · Open weightOpenHands · high reasoningOpenHands
69.80%
2
Grok 4.5xAI · ClosedOpenHands · high reasoningOpenHands
66.03%
3
Gemini 3.7 FlashGoogle · ClosedOpenHands · high reasoningOpenHands
65.89%
4
GPT-5.5 CodexOpenAICodexCodex
62.55%
5
GPT-5.5OpenAI · ClosedOpenHands · xhigh reasoningOpenHands
62.21%
6
Claude Fable 5.1Anthropic · ClosedOpenHandsOpenHands
61.55%
7
GPT-5.6 LunaOpenAI · ClosedOpenHands · max reasoningOpenHands
60.45%
8
Claude Opus 5Anthropic · ClosedOpenHandsOpenHands
60.44%
9
Claude Opus 4.8Anthropic · ClosedOpenHandsOpenHands
59.23%
10
Muse Spark 1.1Meta · ClosedOpenHands · xhigh reasoningOpenHands
59.19%
11
GPT-5.6 TerraOpenAI · ClosedOpenHands · max reasoningOpenHands
58.90%
12
Gemini 3.8 FlashGoogle · ClosedOpenHands · high reasoningOpenHands
57.98%
13
Grok 4.6xAI · ClosedOpenHands · high reasoningOpenHands
55.77%
14
Qwen3.7 PlusAlibaba · ClosedOpenHandsOpenHands
54.30%
15
GPT-5.6 SolOpenAI · ClosedOpenHands · max reasoningOpenHands
54.10%
16
DeepSeek V4 Pro 0813DeepSeek · ClosedOpenHands · max reasoningOpenHands
53.83%
17
Muse Spark 1.2Meta · ClosedOpenHands · xhigh reasoningOpenHands
53.04%
18
Gemini 3.5 FlashGoogle · ClosedOpenHands · high reasoningOpenHands
52.74%
19
GPT-5.4OpenAI · ClosedOpenHands · xhigh reasoningOpenHands
51.71%
20
MiniMax M3MiniMax · Open weightOpenHandsOpenHands
51.50%
21
DeepSeek V4 Pro 0813DeepSeek · ClosedOpenHands · max reasoningOpenHands
51.27%
22
DeepSeek V4 Flash 0731DeepSeek · ClosedOpenHands · high reasoningOpenHands
50.67%
23
Kimi K2.7 CodeMoonshot AI · Open weightOpenHandsOpenHands
50.04%
24
Claude Sonnet 4.6Anthropic · ClosedOpenHandsOpenHands
49.05%
25
GLM-5.3Z.AI · Open weightOpenHands · max reasoningOpenHands
47.51%
26
Claude Sonnet 5Anthropic · ClosedOpenHandsOpenHands
46.48%
27
GLM-5.2Z.AI · Open weightOpenHands · max reasoningOpenHands
45.08%
28
Qwen3.8 MaxAlibaba · Open weightOpenHandsOpenHands
42.01%
29
Grok 4.3xAI · ClosedOpenHands · high reasoningOpenHands
40.64%
30
GLM-5.3-FlashZ.AI · Open weightOpenHands · max reasoningOpenHands
40.16%
31
Qwen3.8-27BAlibaba · Open weightOpenHands · xhigh reasoningOpenHands
38.11%
32
Inkling-SmallThinking Machines Lab · Open weightOpenHands · 0.99 reasoningOpenHands
33.58%
33
Ling 3.0 Flash 2607AntOpenHandsOpenHands
27.95%
34
InklingThinking Machines Lab · Open weightOpenHands · 0.99 reasoningOpenHands
26.09%
35
Mercury 2.5Inception · ClosedOpenHands · high reasoningOpenHands
18.06%

How SkillsBench is shown here

BenchLM mirrors the public Vals AI SkillsBench leaderboard captured from https://www.vals.ai/benchmarks/skillsbench and updated by Vals on September 11, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

SkillsBench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

35 Vals rows3 task viewspublic datasetTasks: Overall, No Skills, With SkillsDisplay only

The published SkillsBench snapshot places DeepSeek V4.1 Flash first at 69.80%. The third row is 3.91 points behind. The broader top-10 range is 10.61 points, so the table still separates the published systems.

35 models have been evaluated on SkillsBench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. SkillsBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About SkillsBench

Year

2026

Tasks

Agent skill-importance tasks

Format

Accuracy score

Difficulty

Agent skill evaluation

BenchLM mirrors the public Vals AI SkillsBench leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

Freshness and provenance

Version

SkillsBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does SkillsBench measure?

How important are skills for agents?

Which model leads the published SkillsBench snapshot?

DeepSeek V4.1 Flash currently leads the published SkillsBench snapshot with 69.80% skillsbench score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on SkillsBench?

The September 11, 2026 snapshot contains 35 AI models.

Last updated: September 11, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.