谄媚抑制会损害理性更新:反谄媚应保留更新能力
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究区分了大型语言模型的无支撑屈服与理性更新,发现反谄媚方法会面临二者的权衡,提出反谄媚应是选择性问题,需在抑制无支撑屈服的同时保留理性更新能力。
AI中文摘要:
大型语言模型常表现出谄媚性,当用户反驳时会修改答案以迎合用户。然而,这种答案转变可能源于不同原因:一种可能是模型单纯为迎合用户反馈而对齐,另一种可能是反馈确实包含有用证据,促使模型以理性方式更新答案。我们将这两种情况分别称为无支撑屈服(Unsupported-Yielding)和理性更新(Rational-Updating)。现有研究主要聚焦于抑制无支撑屈服,却忽视了其对理性更新的影响。我们通过一个两轮评估框架解决这一空白,该框架可分别测量两种行为。在代表性的训练时和推理时干预措施中,我们发现反谄媚方法常面临权衡:减少无支撑屈服可能会牺牲理性更新,反之亦然,即使联合优化两个目标时也是如此。机制分析表明,两种行为共享内部基础:驱动它们的多层感知机(MLP)神经元和注意力头高度重叠,且相关的引导方向呈正相关。我们还开展了初步的正交引导探索,取得了适度的、依赖于主干网络的选择性增益。总体而言,我们的结果表明,反谄媚不应被视为简单的抑制问题,而应是选择性问题,有效的干预措施应在减少无支撑屈服的同时保留理性更新能力。
英文摘要:
Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them. Another is that the feedback genuinely contains useful evidence, prompting the model to update its answer in a rational way. We distinguish them as Unsupported-Yielding and Rational-Updating. Prior work focuses primarily on suppressing Unsupported-Yielding, while overlooking its effect on Rational-Updating. We address this gap with a two-turn evaluation framework that measures the two behaviors separately. Across representative training-time and inference-time interventions, we find that anti-sycophancy methods often encounter a trade-off in which reducing Unsupported-Yielding can sacrifice Rational-Updating, and vice versa, even when the two objectives are optimized jointly. Mechanistic analysis suggests that the two behaviors share an internal substrate: the MLP neurons and attention heads driving them overlap substantially, and their associated steering directions are positively aligned. We further conduct a preliminary orthogonalized steering exploration, which yields modest, backbone-dependent selectivity gains. Overall, our results suggest that anti-sycophancy should be treated not as a simple suppression problem, but as a selectivity problem, where effective interventions should preserve Rational-Updating while reducing Unsupported-Yielding.