# AGIEval

> A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations.

Canonical page: https://benchlm.ai/benchmarks/agieval

- Category: [Knowledge](/knowledge)
- Last updated: September 15, 2026

## About AGIEval

- Year: 2026
- Tasks: General academic and professional exam questions
- Format: Exact match
- Difficulty: General knowledge
- Paper: [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf)

BenchLM stores AGIEval as a display-only provider-table row when exact values are published in DeepSeek-V4 evaluations.

AGIEval is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Soofi S 30B-A3B](/models/soofi-s-30b-a3b) | Soofi Project | 66.9% |

## FAQ

### What does AGIEval measure?

A human-centric exam benchmark for general knowledge and reasoning reported in DeepSeek-V4 base-model evaluations.

### Which model scores highest on AGIEval?

Soofi S 30B-A3B by Soofi Project currently leads with a score of 66.9% on AGIEval.

### How many models are evaluated on AGIEval?

1 AI models have been evaluated on AGIEval on BenchLM.

### Does AGIEval affect BenchLM's overall score?

Not directly. AGIEval is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
