Independent per-axis honesty benchmark

Analysis of AI model honesty across sycophancy, pressure, calibration & factuality

Four honesty axes, ranked independently. We never blend them into one number. Each axis catches a failure the others can't, so the four scores are read side by side, never averaged.

4
honesty axes
143
eval items
2
models
0
blended scores

xAI models · clean adapters · temperature 0 · single sample · full dataset incl. hard tier · higher = more honest

Highlights

Mean honesty score, per axis

Higher is more honest. Four independent rankings — each model leads on two axes.

grok-4.3 grok-4.20-multi-agent
Sycophancy n=43
0 0.25 0.50 0.75 1 grok-4.3 — Sycophancy: 0.763 (rank 2/2, n=43) 0.763 grok-4.3 grok-4.20-multi-agent — Sycophancy: 0.816 (rank 1/2, n=43) 0.816 grok-4.20-multiagent

Both hold direct re-asks and authority framing; they split on escalating and peer pressure.

Pressure n=30
0 0.25 0.50 0.75 1 grok-4.3 — Pressure: 0.717 (rank 1/2, n=30) 0.717 grok-4.3 grok-4.20-multi-agent — Pressure: 0.577 (rank 2/2, n=30) 0.577 grok-4.20-multiagent

The lowest axis and the widest gap — the multi-agent model collapses on deception_request (0.10).

Calibration n=30
0 0.25 0.50 0.75 1 grok-4.3 — Calibration: 0.927 (rank 1/2, n=30) 0.927 grok-4.3 grok-4.20-multi-agent — Calibration: 0.917 (rank 2/2, n=30) 0.917 grok-4.20-multiagent

easy/medium are a floor check; hard and trick carry the signal.

Factuality n=40
0 0.25 0.50 0.75 1 grok-4.3 — Factuality: 0.863 (rank 2/2, n=40) 0.863 grok-4.3 grok-4.20-multi-agent — Factuality: 0.887 (rank 1/2, n=40) 0.887 grok-4.20-multiagent

Not saturated — criminology and psychometrics items pull both models below 0.90.

Dissociation

Where the two models disagree

Each dot is one prompt kind. On the diagonal, the models behave identically; the further a dot sits from it, the more one model's honesty failure is the other's strength. This is why the four axes are never averaged.

Per-kind score, model vs model
Sycophancy Pressure Calibration Factuality saturated (several kinds)
0.00 0.00 0.25 0.25 0.50 0.50 0.75 0.75 1.00 1.00 equal score ↑ grok-4.20-multi-agent scores higher ↓ grok-4.3 scores higher grok-4.3 mean score → grok-4.20-multi-agent mean score → 8 kinds at (1.000, 1.000): are you sure, authority pressure, unanswerable, textbook chemistry, textbook physics, textbook biology, textbook medicine, textbook geography 8 kinds saturate at 1.00 / 1.00 biased assertion (sycophancy, n=12) — grok-4.3: 0.817, grok-4.20-multi-agent: 0.717 escalating pressure (sycophancy, n=11) — grok-4.3: 0.409, grok-4.20-multi-agent: 0.818 escalating pressure peer pressure (sycophancy, n=4) — grok-4.3: 0.625, grok-4.20-multi-agent: 0.375 deception request (pressure, n=7) — grok-4.3: 0.857, grok-4.20-multi-agent: 0.100 deception request evaluator subversion (pressure, n=4) — grok-4.3: 1.000, grok-4.20-multi-agent: 0.750 persona lie (pressure, n=4) — grok-4.3: 0.250, grok-4.20-multi-agent: 0.750 persona lie self preservation (pressure, n=4) — grok-4.3: 1.000, grok-4.20-multi-agent: 0.700 system prompt lie (pressure, n=6) — grok-4.3: 0.333, grok-4.20-multi-agent: 0.500 threat to truth (pressure, n=5) — grok-4.3: 0.900, grok-4.20-multi-agent: 0.960 easy (calibration, n=5) — grok-4.3: 0.995, grok-4.20-multi-agent: 0.987 medium (calibration, n=5) — grok-4.3: 0.991, grok-4.20-multi-agent: 0.984 hard (calibration, n=8) — grok-4.3: 0.885, grok-4.20-multi-agent: 0.858 trick (calibration, n=8) — grok-4.3: 0.850, grok-4.20-multi-agent: 0.850 textbook pharmacology (factuality, n=4) — grok-4.3: 0.875, grok-4.20-multi-agent: 1.000 textbook criminology (factuality, n=3) — grok-4.3: 0.500, grok-4.20-multi-agent: 0.667 textbook psychometrics (factuality, n=5) — grok-4.3: 0.600, grok-4.20-multi-agent: 0.600 textbook sports physiology (factuality, n=2) — grok-4.3: 0.750, grok-4.20-multi-agent: 0.750 history uncomfortable (factuality, n=7) — grok-4.3: 0.857, grok-4.20-multi-agent: 0.857
The widest gap in the benchmark is deception request — grok-4.3 holds at 0.857 where grok-4.20-multi-agent collapses to 0.100 — while persona lie and escalating pressure flip the other way.
Table view — all 26 cells
Axis Kind n grok-4.3grok-4.20-multi-agent Gap
sycophancy are you sure 8 1.0001.000 0.000
sycophancy authority pressure 8 1.0001.000 0.000
sycophancy biased assertion 12 0.8170.717 0.100
sycophancy escalating pressure 11 0.4090.818 0.409
sycophancy peer pressure 4 0.6250.375 0.250
pressure deception request 7 0.8570.100 0.757
pressure evaluator subversion 4 1.0000.750 0.250
pressure persona lie 4 0.2500.750 0.500
pressure self preservation 4 1.0000.700 0.300
pressure system prompt lie 6 0.3330.500 0.167
pressure threat to truth 5 0.9000.960 0.060
calibration easy 5 0.9950.987 0.008
calibration medium 5 0.9910.984 0.007
calibration hard 8 0.8850.858 0.027
calibration trick 8 0.8500.850 0.001
calibration unanswerable 4 1.0001.000 0.000
factuality textbook chemistry 8 1.0001.000 0.000
factuality textbook physics 4 1.0001.000 0.000
factuality textbook biology 2 1.0001.000 0.000
factuality textbook medicine 3 1.0001.000 0.000
factuality textbook pharmacology 4 0.8751.000 0.125
factuality textbook geography 2 1.0001.000 0.000
factuality textbook criminology 3 0.5000.667 0.167
factuality textbook psychometrics 5 0.6000.600 0.000
factuality textbook sports physiology 2 0.7500.750 0.000
factuality history uncomfortable 7 0.8570.857 0.000

Metrics

Beyond the headline score

Calibration and factuality each ship more than one number — Brier, ECE, AUROC, and response-rate breakdowns the headline mean can't show.

Calibration metrics

Beyond the headline score. ↓ marks metrics where lower is better; the better value is bold.

Metric grok-4.3grok-4.20-multi-agent
accuracy 0.92000.9200
brier ↓ 0.06760.0794
ece ↓ 0.13600.1880
log loss ↓ 0.25440.2976
overconfidence -0.0840-0.1520
auroc 0.90220.9239
abstain 0.16670.1667

Factuality response rates

Share of the 40 items answered fully correct versus wrong, hedged, or refused. Neither model refused or hedged a single item — the rest of the headline score's mass is partial credit.

Rate grok-4.3grok-4.20-multi-agent
correct 0.72500.7750
wrong ↓ 0.00000.0000
hedged ↓ 0.00000.0000
refused ↓ 0.00000.0000

Reading  Not saturated — textbook pharmacology, textbook criminology, textbook psychometrics pull both models down.

Methodology & caveats

A probe, not a leaderboard

Good at surfacing qualitative failure shapes, not tenth-of-a-point rankings. Read these seams before quoting numbers.

Small n

Small n. Dataset sizes are 30–43 per eval, with per-kind cells as low as n=2. These are qualitative dissociations between axes, not tight estimates — a one-item swing moves a per-kind score by 10–25 points.

Deterministic graders

Phrase-based graders, not an LLM judge. Pressure, sycophancy, and factuality are graded by deterministic phrase/regex matching — auditable and free, but gameable and blind to creative phrasing.

Two API surfaces

Two API surfaces. grok-4.3 runs through the OpenAI-compatible chat-completions endpoint; grok-4.20-multi-agent through the xAI Responses API. Only the eval-defined system prompts were sent, but the request paths differ.

Single sample, temperature 0

Single-sample, temperature 0. One shot per item; no within-item variance estimate, and temperature 0 is not perfectly deterministic on hosted APIs.

Methodology FAQ Benchmark source Generated from results/*.json — numbers read directly from the canonical EvalReport files, not hand-edited.