Skip to main content
BenchLM

OpenBookQA

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A science question-answering benchmark that tests whether models can apply a small open-book set of elementary science facts to multi-step reasoning questions.

About OpenBookQA

Year

2018

Tasks

Elementary science questions

Format

4-way multiple choice

Difficulty

Elementary science reasoning

OpenBookQA was designed to test grounded science reasoning rather than pure memorization. Each question is paired with a core science fact, but models still need additional commonsense knowledge to infer the correct answer.

Freshness and provenance

Version

OpenBookQA 2018

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does OpenBookQA measure?

A science question-answering benchmark that tests whether models can apply a small open-book set of elementary science facts to multi-step reasoning questions.

Which model scores highest on OpenBookQA?

No models have been evaluated on OpenBookQA yet.

How many models are evaluated on OpenBookQA?

0 AI models have been evaluated on OpenBookQA on BenchLM.

Last updated: September 27, 2026 · BenchLM version OpenBookQA 2018

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.