Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

General AI Assistants (GAIA)

GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.

Data verified 25 confirmed releases in the last 30 daysStart free brief

How BenchLM shows GAIA right now

BenchLM is tracking GAIA in the local dataset, but exact-source verification records for these rows are still being attached. To avoid a blank benchmark page, BenchLM shows the current tracked rows below as a display-only reference table.

These tracked rows are useful for inspection and spot-checking, but until exact-source attachments are completed they should not be treated as fully verified public benchmark rows.

Snapshot

27 tracked modelsLocal tracked rowsAwaiting exact-source attachmentsDisplay only

Tracked score on GAIA — August 21, 2026

We mirror the published tracked score view for GAIA. Claude Mythos 5 leads the public snapshot at 52.3%, followed by Claude Fable 5 (52.3%) and GPT-5.4 Pro (50.5%). We do not use these results to rank models overall.

27 modelsAgenticRefreshingDisplay onlyUpdated August 21, 2026

Tracked score table (27 models)

Score
1
Claude Mythos 5Anthropic · Closed
52.3%
2
Claude Fable 5Anthropic · Closed
52.3%
3
GPT-5.4 ProOpenAI · Closed
50.5%
4
GPT-5.4OpenAI · Closed
48.2%
5
Claude Opus 4.6Anthropic · Closed
47.8%
6
Gemini 3.1 ProGoogle · Closed
46.1%
7
Claude Sonnet 4.6Anthropic · Closed
45.5%
8
GPT-5 miniOpenAI · Closed
44.8%
9
GPT-5 (high)OpenAI · Closed
42.1%
10
GPT-5.2OpenAI · Closed
40.3%
11
Grok 4.1xAI · Closed
39.7%
12
Gemini 3 ProGoogle · Closed
38.5%
13
Kimi K2.5 (Reasoning)Moonshot AI · Closed
38.1%
14
Qwen3.6 PlusAlibaba · Closed
37.4%
15
o4-mini (high)OpenAI · Closed
36.8%
16
Kimi K2.5Moonshot AI · Open weight
36.5%
17
Qwen3.5 397B (Reasoning)Alibaba · Open weight
35.6%
18
Gemini 3 FlashGoogle · Closed
35.2%
19
DeepSeek V3.2 (Thinking)DeepSeek · Open weight
34.8%
20
Llama 4 BehemothMeta · Open weight
34.2%
21
GLM-5 (Reasoning)Z.AI · Open weight
33.8%
22
Qwen3.5 397BAlibaba · Open weight
33.2%
23
DeepSeek V3.2DeepSeek · Open weight
32.4%
24
GLM-5Z.AI · Open weight
31.5%
25
Mistral Large 3Mistral · Closed
30.8%
26
Step 3.5 FlashStepFun · Open weight
29.4%
27
Llama 4 MaverickMeta · Open weight
28.6%

The published GAIA snapshot places Claude Mythos 5 first at 52.3%. The third row is 1.8 points behind. The broader top-10 range is 12.0 points, so the table still separates the published systems.

27 models have been evaluated on GAIA. The benchmark falls in the Agentic category. This category carries a 22% weight in BenchLM.ai's overall scoring system. GAIA is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About GAIA

Year

2024

Tasks

466

BenchLM freshness & provenance

Version

GAIA 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does GAIA measure?

GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.

Which model leads the published GAIA snapshot?

Claude Mythos 5 currently leads the published GAIA snapshot with 52.3% tracked score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on GAIA?

27 AI models are included in BenchLM's mirrored GAIA snapshot, based on the public leaderboard captured on August 21, 2026.

Last updated: August 21, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.