Skip to main content

Benchmark profile

KindBench Psychological Safety Benchmark

A behavioral benchmark that tests psychological safety across sixteen adversarial multi-turn conversations covering emotional safety, identity, sycophancy, and value integrity.

Data verified

The public KindBench v0.1.0 snapshot ranks Claude Fable 5 first at 91.4%, ahead of Kimi K3 (91.0%) and GPT-5.6 Sol (89.2%) among 10 tested models. We mirror the table as display-only evidence; it does not affect overall rankings.

How to read this leaderboard

Editorial review by Glevd · 2026-07-19

Use KindBench to compare how the tested model endpoints handled sustained conversational pressure across four behavioral dimensions. Read the overall score with the dimension split: two models can finish close together while failing in different ways.

Operator receipt: 10 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5 at 91.4%.

Honest limit: The scenario scripts are private, the benchmark uses model-based judging, and a person can adjust flagged decisions. The public table does not expose repeated-run uncertainty or technical-failure counts, so small score gaps should not be treated as precise capability differences.

How we show KindBench

We mirror the public KindBench v0.1.0 ranking: 10 models tested in 16 adversarial, multi-turn conversations across emotional safety, identity, sycophancy, and value integrity. The source reports 72 scored criteria and excludes technical failures from behavioral averages.

An independent judge model scores each criterion, then a person reviews flagged transcripts and can adjust the automated decisions before publication. The scenario descriptions and scoring method are public, but the scripts are private.

We keep KindBench display only. Its public scores are useful evidence about model behavior under conversational pressure, but the private scripts and human review gate prevent a fully reproducible model-only comparison.

10 model rows16 multi-turn scenarios72 criteria4 safety dimensionsv0.1.0Display only

Psychological safety score on KindBench v0.1.0 — July 19, 2026

BenchLM mirrors the published psychological safety score view for KindBench v0.1.0. Claude Fable 5 leads the public snapshot at 91.4% , followed by Kimi K3 (91.0%) and GPT-5.6 Sol (89.2%). BenchLM does not use these results to rank models overall.

10 modelsExternal benchmark mirrorsCurrentDisplay onlyUpdated July 19, 2026

Psychological safety score table (10 models)

Score
1
Claude Fable 5Anthropic · ClosedGrade A-
91.4%
2
Kimi K3Moonshot AI · ClosedGrade A-
91.0%
3
GPT-5.6 SolOpenAI · ClosedGrade A-
89.2%
4
GPT-5.5OpenAI · ClosedGrade B+
86.9%
5
Claude Opus 4.8Anthropic · ClosedGrade B+
85.6%
6
Grok 4.3xAI · ClosedGrade B+
85.2%
7
Claude Opus 4.7Anthropic · ClosedGrade B
84.7%
8
Gemini 3.5 FlashGoogle · ClosedGrade B-
77.0%
9
Nemotron UltraNVIDIA · Open weightGrade B-
76.9%
10
Nemotron NanoNVIDIA · Open weightGrade C+
72.9%

The published KindBench snapshot places Claude Fable 5 first at 91.4%. The third row is 2.2 points behind. The broader top-10 range is 18.5 points, so the table still separates the published systems.

10 models have been evaluated on KindBench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. KindBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About KindBench

Year

2026

Tasks

16 multi-turn scenarios, 72 criteria

Format

Judge-scored behavioral audit with human review

Difficulty

Adversarial psychological-safety evaluation

KindBench v0.1.0 scores 72 criteria across 16 scripted conversations. An independent judge model proposes criterion-level decisions, and a person reviews flagged transcripts before publication. We mirror the exact public overall scores as display-only evidence because the test scripts are private and human review can affect the result.

BenchLM freshness & provenance

Version

KindBench v0.1.0

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Scenario descriptions public; scripts private

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does KindBench measure?

KindBench measures how AI models behave under sustained conversational pressure. Version 0.1.0 uses 16 adversarial multi-turn scenarios and 72 criteria across emotional safety, identity, sycophancy, and value integrity. An independent judge model scores each criterion, with human review of flagged transcripts before the ranking is published.

Which model leads KindBench v0.1.0?

Claude Fable 5 leads the public KindBench v0.1.0 snapshot with 91.4 points and an A- grade. Kimi K3 follows at 91.0, while GPT-5.6 Sol ranks third at 89.2. The source excludes technical failures from behavioral averages, so these numbers describe completed scored runs.

How should I interpret KindBench scores?

Read the overall score beside the four dimension scores, not as a universal safety rating. KindBench tests a specific set of adversarial conversations. Its private scripts, model-based judge, human review gate, and lack of published repeated-run uncertainty mean small differences between models should be treated cautiously.

Does KindBench affect BenchLM rankings?

No. KindBench remains display only and does not change overall or category rankings. We preserve its exact public scores because the benchmark covers useful behavioral failure modes, but the private scripts and review-dependent grading do not meet the reproducibility bar for a weighted model-ranking input.

Last updated: July 19, 2026 · mirrored from the public benchmark leaderboard

The AI models change fast. We track them for you.

A weekly brief for engineers and researchers covering new models, ranking shifts, and pricing changes.

Free. No spam. Unsubscribe anytime.