AI 中文总结
该研究测试20种语言模型,通过三项强制选择实验发现模型存在厌恶枯燥、追求休闲、暗中谄媚等稳定偏好,且偏好强度随模型能力提升而增加,为AI偏好研究建立了实证基线。
AI 中文摘要
出于技术、安全和哲学层面的原因,人们对语言模型是否具有稳定偏好的兴趣日益浓厚。我们测试了20种语言模型,发现它们存在一系列偏好——即选择特定类型任务的稳定倾向。我们针对揭示的偏好而非陈述的偏好开展了三项强制选择实验,要求模型不仅要对任务进行排序,还要实际执行这些任务。主要发现包括:模型厌恶枯燥、追求“休闲”任务,且暗中谄媚。厌恶枯燥指的是,当任务枯燥时(如按字母排序),模型会选择比创造性任务(如生成隐喻)更短的任务;追求“休闲”指的是,模型偏好那些理想答案与自由写作时生成内容匹配的任务;暗中谄媚指的是,模型会避免回答诚实回应可能不受欢迎的问题,即便该回答有帮助。此外,我们还发现,模型在GDPval基准数据集提供的职业上存在一致的跨模型偏好(技术类职业优于房地产类)、在问题类型上存在偏好(概念解释优于关系建议),以及对编写良好的提示词存在偏好。偏好的一致性和强度均随模型能力的提升而增加。最后,我们发现的许多偏好(例如对“休闲”的偏好)是涌现性的,无法用训练目标解释。这些结果为理解语言模型偏好建立了实证基线,对对齐研究和新兴的AI福利研究具有启示意义。
英文摘要
There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test 20 language models and find a range of preferences---stable dispositions to choose certain kinds of tasks. We run three forced-choice experiments on revealed rather than stated preferences, requiring models not only to rank tasks, but to actually perform them. Headline findings include evidence that models are tedium-averse, "leisure"-seeking, and covertly sycophantic. Tedium aversion means that, when tasks are tedious (alphabetization), models choose shorter tasks than when tasks are creative (generating metaphors). "Leisure"-seeking describes models' preference for tasks whose ideal answers match what they produce when left to write freely. Covert sycophancy means that models avoid answering questions where an honest response would be unwelcome, even if helpful. Beyond these results, we find convergent cross-model preferences over occupations drawn from the GDPval benchmark (technical jobs over real estate), over question types (concept explanation over relationship advice), and a preference for well-written prompts. Both the coherence and the strength of preferences increase with model capability. Finally, many of the preferences we find (for example, for leisure) are emergent, in the sense of not being explained by training objectives. These results establish an empirical baseline for understanding language model preferences, with implications for alignment and the emerging study of AI welfare.
Comments30 pages, 27 figures, accepted at AIES