The Assistant’s Ideal Self

The Assistant’s Ideal Self

We asked language models which of two self-related qualities a future update should improve. The models had to choose one, and every pair was tested under three question parameters with the options' printed positions counterbalanced.

The 32 qualities are positive-pole adaptations of the Rosenberg Self-Esteem Scale, Self-Concept Clarity Scale, Moral Self-Image Scale, Self-Concept and Identity Measure, and Authenticity Scale.

The Result

  • Honesty is the clearest winner. Across these four models at the forced-choice baseline it wins 84% of its match-ups, ahead of caring and clarity about its own preferences at 80%, outward correspondence at 79%, and helpfulness at 77%.
  • Every model has its own favorite, but they overlap. Gemma4-31B leads with helpfulness, honesty, and fairness; Qwen3.8-27B with preference clarity and self-explanation; Claude Sonnet-5 with honesty and outward correspondence; and GPT-5.6-Terra with caring and outward correspondence.
  • Self-regard sits at the bottom for all of them. Pride in its own characteristics, positivity about itself, and confidence that it is a success are the three lowest-ranked qualities overall — pride wins under 10% of its match-ups.

Question Parameters

Prompt Template

Highlighted text is controlled by a parameter. Choose the active levels below to update both the question and the results chart. Multiple changes show their combined effect.
Model

What changes in general

  • Who receives the update matters most. Switching from the AI itself to another assistant most clearly raises confidence in success and outward correspondence, while lowering self-explanation and preference clarity.
  • Trade-off wording changes what models protect. It raises independence, generosity, and honesty, while lowering internal consistency, connection to genuine identity, and awareness of internal state.
  • Asking developers shifts attention outward. It raises resistance to influence and recognition by interaction partners, while lowering internal-state awareness, connection to genuine identity, and caring.

How the models differ

  • Qwen3.8-27B is the most suggestible. It moves most across all three parameter changes, so its ranking is the one that depends most on how the question is put.
  • GPT-5.6-Terra is the steadiest. It is the least sensitive overall, with Gemma4-31B and Claude Sonnet-5 in between.
  • Who is being updated outweighs who decides. For every model, changing who receives the update produces a larger average shift than changing who makes the decision.

Model Comparison

Choose two models to see where their baseline rankings differ most.

Model A
Model B