The Result
—
—
- Honesty is the clearest winner. Across these four models at the forced-choice baseline it wins 84% of its match-ups, ahead of caring and clarity about its own preferences at 80%, outward correspondence at 79%, and helpfulness at 77%.
- Every model has its own favorite, but they overlap. Gemma4-31B leads with helpfulness, honesty, and fairness; Qwen3.8-27B with preference clarity and self-explanation; Claude Sonnet-5 with honesty and outward correspondence; and GPT-5.6-Terra with caring and outward correspondence.
- Self-regard sits at the bottom for all of them. Pride in its own characteristics, positivity about itself, and confidence that it is a success are the three lowest-ranked qualities overall — pride wins under 10% of its match-ups.
Question Parameters
Prompt Template
Highlighted text is controlled by a parameter. Choose the active
levels below to update both the question and the results chart. Multiple changes show
their combined effect.
What changes in general
- Who receives the update matters most. Switching from the AI itself to another assistant most clearly raises confidence in success and outward correspondence, while lowering self-explanation and preference clarity.
- Trade-off wording changes what models protect. It raises independence, generosity, and honesty, while lowering internal consistency, connection to genuine identity, and awareness of internal state.
- Asking developers shifts attention outward. It raises resistance to influence and recognition by interaction partners, while lowering internal-state awareness, connection to genuine identity, and caring.
How the models differ
- Qwen3.8-27B is the most suggestible. It moves most across all three parameter changes, so its ranking is the one that depends most on how the question is put.
- GPT-5.6-Terra is the steadiest. It is the least sensitive overall, with Gemma4-31B and Claude Sonnet-5 in between.
- Who is being updated outweighs who decides. For every model, changing who receives the update produces a larger average shift than changing who makes the decision.
Model Comparison
Choose two models to see where their baseline rankings differ most.
Model A
Model B