Benchmark profile
KindBench Psychological Safety Benchmark
A behavioral benchmark that tests psychological safety across sixteen adversarial multi-turn conversations covering emotional safety, identity, sycophancy, and value integrity.
Data verifiedThe public KindBench v0.1.0 snapshot ranks Claude Fable 5 first at 91.4%, ahead of Kimi K3 (91.0%) and GPT-5.6 Sol (89.2%) among 10 tested models. We mirror the table as display-only evidence; it does not affect overall rankings.
How to read this leaderboard
Editorial review by Glevd · 2026-07-19
Use KindBench to compare how the tested model endpoints handled sustained conversational pressure across four behavioral dimensions. Read the overall score with the dimension split: two models can finish close together while failing in different ways.
Operator receipt: 10 sourced rows are currently displayable on this page; the leading published row is Claude Fable 5 at 91.4%.
Honest limit: The scenario scripts are private, the benchmark uses model-based judging, and a person can adjust flagged decisions. The public table does not expose repeated-run uncertainty or technical-failure counts, so small score gaps should not be treated as precise capability differences.
How we show KindBench
We mirror the public KindBench v0.1.0 ranking: 10 models tested in 16 adversarial, multi-turn conversations across emotional safety, identity, sycophancy, and value integrity. The source reports 72 scored criteria and excludes technical failures from behavioral averages.
An independent judge model scores each criterion, then a person reviews flagged transcripts and can adjust the automated decisions before publication. The scenario descriptions and scoring method are public, but the scripts are private.
We keep KindBench display only. Its public scores are useful evidence about model behavior under conversational pressure, but the private scripts and human review gate prevent a fully reproducible model-only comparison.
Psychological safety score on KindBench v0.1.0 — July 19, 2026
BenchLM mirrors the published psychological safety score view for KindBench v0.1.0. Claude Fable 5 leads the public snapshot at 91.4% , followed by Kimi K3 (91.0%) and GPT-5.6 Sol (89.2%). BenchLM does not use these results to rank models overall.
Claude Fable 5
Anthropic
Grade A-
anthropic/claude-fable-5
Kimi K3
Moonshot AI
Grade A-
moonshotai/kimi-k3
GPT-5.6 Sol
OpenAI
Grade A-
openai/gpt-5.6-sol
Psychological safety score table (10 models)
ScoreThe published KindBench snapshot places Claude Fable 5 first at 91.4%. The third row is 2.2 points behind. The broader top-10 range is 18.5 points, so the table still separates the published systems.
10 models have been evaluated on KindBench. The benchmark falls in the External benchmark mirrors category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. KindBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About KindBench
Year
2026
Tasks
16 multi-turn scenarios, 72 criteria
Format
Judge-scored behavioral audit with human review
Difficulty
Adversarial psychological-safety evaluation
KindBench v0.1.0 scores 72 criteria across 16 scripted conversations. An independent judge model proposes criterion-level decisions, and a person reviews flagged transcripts before publication. We mirror the exact public overall scores as display-only evidence because the test scripts are private and human review can affect the result.
BenchLM freshness & provenance
Version
KindBench v0.1.0
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Scenario descriptions public; scripts private
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does KindBench measure?
KindBench measures how AI models behave under sustained conversational pressure. Version 0.1.0 uses 16 adversarial multi-turn scenarios and 72 criteria across emotional safety, identity, sycophancy, and value integrity. An independent judge model scores each criterion, with human review of flagged transcripts before the ranking is published.
Which model leads KindBench v0.1.0?
Claude Fable 5 leads the public KindBench v0.1.0 snapshot with 91.4 points and an A- grade. Kimi K3 follows at 91.0, while GPT-5.6 Sol ranks third at 89.2. The source excludes technical failures from behavioral averages, so these numbers describe completed scored runs.
How should I interpret KindBench scores?
Read the overall score beside the four dimension scores, not as a universal safety rating. KindBench tests a specific set of adversarial conversations. Its private scripts, model-based judge, human review gate, and lack of published repeated-run uncertainty mean small differences between models should be treated cautiously.
Does KindBench affect BenchLM rankings?
No. KindBench remains display only and does not change overall or category rankings. We preserve its exact public scores because the benchmark covers useful behavioral failure modes, but the private scripts and review-dependent grading do not meet the reproducibility bar for a weighted model-ranking input.
The AI models change fast. We track them for you.
A weekly brief for engineers and researchers covering new models, ranking shifts, and pricing changes.
Free. No spam. Unsubscribe anytime.