# A Benchmark for Visual Question Answering using World Knowledge (A-OKVQA)

> A visual question answering benchmark whose questions cannot be answered from the image alone and require commonsense or world knowledge.

Canonical page: https://benchlm.ai/benchmarks/aokvqa

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 27, 2026

## About A-OKVQA

- Year: 2022
- Tasks: Knowledge-grounded visual question answering
- Format: Multiple choice and direct answer
- Difficulty: Commonsense and world knowledge about images
- Paper: [A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge](https://arxiv.org/abs/2206.01718)

BenchLM stores A-OKVQA as a display-only knowledge-grounded visual QA reference outside the weighted core schema. Providers report either the multiple-choice or the direct-answer setting, so the row note should record which split and setting produced the value.

A-OKVQA is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Ternary Bonsai 2 27B](/models/ternary-bonsai-2-27b) | Prism ML | 86.8% |

## FAQ

### What does A-OKVQA measure?

A visual question answering benchmark whose questions cannot be answered from the image alone and require commonsense or world knowledge.

### Which model scores highest on A-OKVQA?

Ternary Bonsai 2 27B by Prism ML currently leads with a score of 86.8% on A-OKVQA.

### How many models are evaluated on A-OKVQA?

1 AI models have been evaluated on A-OKVQA on BenchLM.

### Does A-OKVQA affect BenchLM's overall score?

Not directly. A-OKVQA is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
