接受性,而非谄媚:区分语言模型中的参与与顺从
Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models
- Harvard Kennedy School, Harvard University(哈佛大学肯尼迪学院)
- Department of Statistics, Harvard University(哈佛大学统计系)
- Computer Science, Stanford University(斯坦福大学计算机科学系)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究指出社会性谄媚评估与对话接受性重叠导致构念效度问题,通过实验证明接受性更受青睐,并提出方法在不增加顺从的同时提升接受性。
AI中文摘要:
语言模型的一个核心关注点是谄媚:它们倾向于顺从用户的观点,而牺牲独立的实质性判断。与此同时,关于社会性谄媚的研究聚焦于诸如认可和积极性等可能暗示不当顺从的行为。然而,社会性谄媚的标志也是对话接受性的特征,后者是社会心理学中的一个构念,已被证明能改善分歧中的互动。我们认为,这种重叠为社会性谄媚评估带来了构念效度问题。使用一个流行的道德建议数据集,我们发现被归类为更具社会性谄媚的回应也更具接受性。此外,提高人工撰写的回应的接受性——同时保留其实质性结论——会导致它们被归类为更具社会性谄媚。这种紧密耦合提出了可能性,即社会性谄媚评估无意中惩罚了可取的行为。在一项预先注册的实验中,比较实质等效的回应,参与者更喜欢更具接受性的回应,期望用户更有可能倾听它们,并且更愿意向其作者寻求建议。即使在认为原始提问者错误的参与者中,同样的总体模式也持续存在。最后,我们引入了一种简单的方法,在不增加实质性顺从的情况下显著提高接受性,证明对话接受性和实质性独立可以同时实现。
英文摘要:
A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment. In parallel, work on social sycophancy has focused on behaviors such as validation and positivity that may signal inappropriate deference. Yet the markers of social sycophancy are also characteristic of conversational receptiveness, a construct from social psychology shown to improve interactions across disagreement. We argue that this overlap creates a construct-validity problem for social sycophancy evaluations. Using a popular moral-advice dataset, we find that responses classified as more socially sycophantic are also more receptive. Further, increasing the receptiveness of human-written responses---while preserving their substantive conclusions---causes them to be classified as more socially sycophantic. This tight coupling raises the possibility that social sycophancy evaluations inadvertently penalize desirable behavior. In a preregistered experiment comparing substantively equivalent responses, participants prefer the more receptive responses, expect users to be more likely to listen to them, and are more willing to seek advice from their authors. The same overall pattern persists even among participants who believe the original question asker is in the wrong. Finally, we introduce a simple approach that substantially increases receptiveness without increasing substantive deference, demonstrating that conversational receptiveness and substantive independence can be achieved together.