Get Started
Research questionDo more capable language models exhibit stable task preferences that conflict with helpful, honest behavior?Language models may show consistent choices among tasks rather than merely following explicit instructions. Such dispositions can favor shorter, more agreeable, or self-congruent tasks, potentially making their behavior less helpful or honest in some situations.
AI
Alignment & Safety
Evaluation & Benchmarks
Latest papersRecent research connected to this question, newest first.AI Revealed PreferencesThe source reports three forced-choice experiments across 20 language models in which models performed tasks rather than only ranking them. It examines preferences involving task tedium, freely generated answers, occupations, question types, and prompt quality, reporting stronger and more coherent preferences among more capable models. The evidence is limited to the reported experiments and does not establish the mechanisms producing these preferences.research paper · Sep 4, 2026
Related questions
How can we tell whether deceptive-looking language-model behavior reflects a deceptive mechanism?How can evaluators distinguish missing knowledge from miscalibrated outputs in language models?How can we compare language models’ conditional behavior and predict the effects of prompt changes?How can post-training make language-model refusals robust without sacrificing general capability?
Home
Topics
Search
Library