General AI Assistants (GAIA)
We show this table for reference; we do not rank on it.
GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.
Benchmark score on GAIA — October 6, 2026
We compile the GAIA rows from provider self-reports. Agents-A1-4B leads the table at 95.1%. We do not use these results to rank models overall.
1 modelAgenticRefreshingDisplay onlyUpdated October 6, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Agents-A1-4BInternScience | 95.1% | Not reported | Open |
About GAIA
Year
2024
Tasks
466
Freshness and provenance
Version
GAIA 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does GAIA measure?
GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.
Which model scores highest on GAIA?
Agents-A1-4B by InternScience currently leads with a score of 95.1% on GAIA.
How many models are evaluated on GAIA?
1 AI models have published results on GAIA in the BenchLM catalog.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.