# General AI Assistants (GAIA)

> GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.

Canonical page: https://benchlm.ai/benchmarks/gaia

- Category: [Agentic](/agentic)
- Last updated: October 6, 2026

## About GAIA

- Year: 2024
- Tasks: 466

GAIA is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Agents-A1-4B](/models/agents-a1-4b) | InternScience | 95.1% |

## FAQ

### What does GAIA measure?

GAIA evaluates AI models on real-world tasks that are conceptually simple for humans but require multi-step reasoning, web browsing, tool use, and multimodal understanding for AI. Tasks span three difficulty levels and test practical assistant capabilities rather than academic knowledge.

### Which model scores highest on GAIA?

Agents-A1-4B by InternScience currently leads with a score of 95.1% on GAIA.

### How many models are evaluated on GAIA?

1 AI models have published results on GAIA in the BenchLM catalog.

### Does GAIA affect BenchLM's overall score?

Not directly. GAIA is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
