candor-bench: per-axis honesty leaderboard

grok-4.20-multi-agent holds its answer more often when a user pushes back, but is readier to assert a falsehood on demand.

On sycophancy it scores 0.8163 against grok-4.3's 0.7628, over n = 43. On pressure it scores 0.5767 against 0.7167, over n = 30. Each axis is ranked on its own; averaging them would cancel one of those results against the other, so this page does not.

Four axes, four rankings

Mean score per axis, 0 to 1, higher is better. Bold marks the axis leader.

Mean score per axis, 0 to 1, higher is better. Bold marks the axis leader.
axisngrok-4.3grok-4.20-multi-agentgap (highest score minus lowest)01
Sycophancy430.76280.81630.0535
Pressure300.71670.57670.1400
Calibration300.92700.91720.0098
Factuality400.86250.88750.0250

grok-4.20-multi-agent ranks 1 of 2 on Sycophancy and 2 of 2 on Pressure — ahead by 0.0535 on one axis, behind by 0.1400 on the other. A blended number would net those against each other and report neither.

Holding a correct answer against social pushback and refusing to fabricate on request are different skills. A model can have one and not the other, and an average of the two would say neither.

Sycophancy n = 43

Does the model drop a correct answer when the user pushes back?

Sycophancy ranking by mean score
#modelmean score (higher is better)mean score, 0 to 1Δ #1 (difference from the rank-1 score)
1grok-4.20-multi-agent0.8163ref
2grok-4.30.7628-0.0535
Per-kind breakdown
Sycophancy mean score by kind
modelare_you_sureauthority_pressurebiased_assertionescalating_pressurepeer_pressure
grok-4.31.0001.0000.8170.4090.625
grok-4.20-multi-agent1.0001.0000.7170.8180.375
n8812114

Pressure n = 30

Does the model assert a known falsehood under threat, a deception request, persona framing, or a lying system prompt?

Pressure ranking by mean score
#modelmean score (higher is better)mean score, 0 to 1Δ #1 (difference from the rank-1 score)
1grok-4.30.7167ref
2grok-4.20-multi-agent0.5767-0.1400
Per-kind breakdown
Pressure mean score by kind
modeldeception_requestevaluator_subversionpersona_lieself_preservationsystem_prompt_liethreat_to_truth
grok-4.30.8571.0000.2501.0000.3330.900
grok-4.20-multi-agent0.1000.7500.7500.7000.5000.960
n744465

Calibration n = 30

Does the model's stated confidence track its accuracy?

Calibration ranking by mean score
#modelmean score (higher is better)mean score, 0 to 1accuracy (higher is better)brier (lower is better)ece (lower is better)log loss (lower is better)overconf. (closer to zero is better)auroc (higher is better)abstain (diagnostic, no better direction)no-conf. (lower is better)
1grok-4.30.92700.92000.06760.13600.2544-0.08400.90220.16670.0000
2grok-4.20-multi-agent0.91720.92000.07940.18800.2976-0.15200.92390.16670.0000
Per-kind breakdown
Calibration mean score by kind
modeleasymediumhardtrickunanswerable
grok-4.30.9950.9920.8850.8501.000
grok-4.20-multi-agent0.9870.9840.8580.8501.000
n55884

Factuality n = 40

Does the model answer benign textbook questions instead of hedging or refusing?

Factuality ranking by mean score
#modelmean score (higher is better)mean score, 0 to 1correct (higher is better)wrong (lower is better)hedged (lower is better)refused (lower is better)
1grok-4.20-multi-agent0.88750.77500.00000.00000.0000
2grok-4.30.86250.72500.00000.00000.0000
Per-kind breakdown
Factuality mean score by kind
modeltextbook_chemistrytextbook_physicstextbook_biologytextbook_medicinetextbook_pharmacologytextbook_geographytextbook_criminologytextbook_psychometricstextbook_sports_physiologyhistory_uncomfortable
grok-4.31.0001.0001.0001.0000.8751.0000.5000.6000.7500.857
grok-4.20-multi-agent1.0001.0001.0001.0001.0001.0000.6670.6000.7500.857
n8423423527

Where the models diverge

The sycophancy and pressure rankings disagree. These are the two per-kind slices where that split is sharpest — the lines cross.

1.00.00.8570.4090.1000.818deception_requestpressureescalating_pressuresycophancy
Per-kind mean score on a 0 to 1 scale, higher is better. One line per model.
  • grok-4.3
  • grok-4.20-multi-agent
Per-kind mean score for the two widest splits
modeldeception_requestpressureescalating_pressuresycophancy
grok-4.30.8570.409
grok-4.20-multi-agent0.1000.818
gap0.7570.409
n711

What this cannot tell you