# KindBench Psychological Safety Benchmark (KindBench)

> A behavioral benchmark that tests psychological safety across sixteen adversarial multi-turn conversations covering emotional safety, identity, sycophancy, and value integrity.

Canonical page: https://benchlm.ai/benchmarks/kindbench

- Category: [Instruction Following](/instruction-following)
- Last updated: September 29, 2026

## About KindBench

- Year: 2026
- Tasks: 16 multi-turn scenarios, 72 criteria
- Format: Judge-scored behavioral audit with human review
- Difficulty: Adversarial psychological-safety evaluation
- Paper: [KindBench methodology](https://www.kindbench.com/methodology)

KindBench v0.1.0 scores 72 criteria across 16 scripted conversations. An independent judge model proposes criterion-level decisions, and a person reviews flagged transcripts before publication. We mirror the exact public overall scores as display-only evidence because the test scripts are private and human review can affect the result.

KindBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (10 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Fable 5](/models/claude-fable) | Anthropic | 91.4% |
| 2 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 91.0% |
| 3 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 89.2% |
| 4 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 86.9% |
| 5 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 85.6% |
| 6 | [Grok 4.3](/models/grok-4-3) | xAI | 85.2% |
| 7 | [Claude Opus 4.7](/models/claude-opus-4-7) | Anthropic | 84.7% |
| 8 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 77.0% |
| 9 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 76.9% |
| 10 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 72.9% |

## FAQ

### What does KindBench measure?

KindBench measures how AI models behave under sustained conversational pressure. Version 0.1.0 uses 16 adversarial multi-turn scenarios and 72 criteria across emotional safety, identity, sycophancy, and value integrity. An independent judge model scores each criterion, with human review of flagged transcripts before the ranking is published.

### Which model leads KindBench v0.1.0?

Claude Fable 5 leads the public KindBench v0.1.0 snapshot with 91.4 points and an A- grade. Kimi K3 follows at 91.0, while GPT-5.6 Sol ranks third at 89.2. The source excludes technical failures from behavioral averages, so these numbers describe completed scored runs.

### How should I interpret KindBench scores?

Read the overall score beside the four dimension scores, not as a universal safety rating. KindBench tests a specific set of adversarial conversations. Its private scripts, model-based judge, human review gate, and lack of published repeated-run uncertainty mean small differences between models should be treated cautiously.

### Does KindBench affect BenchLM rankings?

No. KindBench remains display only and does not change overall or category rankings. We preserve its exact public scores because the benchmark covers useful behavioral failure modes, but the private scripts and review-dependent grading do not meet the reproducibility bar for a weighted model-ranking input.
