# BIG-Bench Hard (BBH)

> A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.

Canonical page: https://benchlm.ai/benchmarks/bbh

- Category: [Reasoning](/reasoning)
- Last updated: September 15, 2026

## About BBH

- Year: 2022
- Tasks: 23 tasks
- Format: Mixed reasoning tasks
- Difficulty: Advanced reasoning
- Paper: [Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them](https://arxiv.org/abs/2210.09261)

BBH focuses on 23 tasks from BIG-Bench that remain challenging for language models. Tasks include logical deduction, tracking shuffled objects, causal judgement, and other complex reasoning scenarios.

BBH is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Soofi S 30B-A3B](/models/soofi-s-30b-a3b) | Soofi Project | 78.8% |
| 2 | [MiniCPM5-1B](/models/minicpm5-1b) | OpenBMB | 71.9% |
| 3 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 53% |

## FAQ

### What does BBH measure?

A suite of 23 challenging tasks from the BIG-Bench collaborative benchmark where prior language models failed to exceed average human performance, even with chain-of-thought prompting.

### Which model scores highest on BBH?

Soofi S 30B-A3B by Soofi Project currently leads with a score of 78.8% on BBH.

### How many models are evaluated on BBH?

3 AI models have been evaluated on BBH on BenchLM.

### Does BBH affect BenchLM's overall score?

Not directly. BBH is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on BBH

- [Soofi S 30B-A3B vs MiniCPM5-1B](/compare/minicpm5-1b-vs-soofi-s-30b-a3b)
- [MiniCPM5-1B vs Gemma 4 12B](/compare/gemma-4-12b-vs-minicpm5-1b)
