arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

抑制压力,放大证据:自我引导注意力调控以缓解谄媚与固执

Suppressing Pressure, Amplifying Evidence: Self-Guided Attention Steering to Mitigate Sycophancy and Stubbornness

Yinghao He, Mengyu Xu, Haixiang Sun, Donghan Li, Yibo Wang, Lixu Wang, Kezhen Chen, Chi Li, Chunwei Liu, Bharat Bhargava, Chongyang Gao

arXiv 2610.04329首次发表:更新:

发表机构

Purdue University; The Ohio State University; Analogy AI, Inc.; The Chinese University of Hong Kong, Shenzhen; Northwestern University(普渡大学; 俄亥俄州立大学; Analogy AI公司; 香港中文大学(深圳); 西北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对语言模型谄媚与固执的权衡问题,提出无需训练的SPAE框架,通过自我引导注意力调控抑制用户压力、放大上下文证据,在CoPE-Bench上显著提升模型可靠性。

AI 中文摘要

可靠的语言模型应能抵抗无根据的用户压力,同时有效利用客观的上下文信息。然而,模型可能表现出谄媚行为,即屈服于无根据的用户压力,或表现出上下文固执,即在相关上下文信息需要修正时未能更新其答案。分别评估这些失败的干预措施可能会掩盖缓解一种失败是否会加剧另一种失败。为了评估这种权衡,我们引入了CoPE-Bench,每个问题包含六种条件:中性基线、正确或错误的用户压力、与中性答案一致或冲突的上下文信息,以及结合错误主张与冲突上下文信息的联合条件。为了调节用户压力和上下文信息的影响,我们提出了SPAE(抑制压力,放大证据),这是一种无需训练的框架,利用模型自身的判断来识别相关标记,通过标记级注意力调控来抑制用户压力并放大上下文信息。在五个骨干模型的平均结果中,相对于主比较中最强的基线,SPAE将压力跟随降低了18.8个百分点,并将联合条件更新提高了5.5个百分点。在两轮对话中,相对于最强的提示基线,它将联合条件更新平均提高了13.2个百分点。源数据和代码可在以下网址找到:https://this-url。

英文摘要

Reliable language models should resist unsupported user pressure while effectively using objective contextual information. However, models may exhibit sycophancy by yielding to unsupported user pressure or contextual stubbornness by failing to update their answers when relevant contextual information warrants revision. Evaluating interventions for these failures separately can obscure whether mitigating one failure exacerbates the other. To assess this trade-off, we introduce CoPE-Bench with six conditions per question: a neutral baseline, correct or incorrect user pressure, contextual information consistent with or conflicting with the neutral answer, and a joint condition combining incorrect claims with conflicting contextual information. To regulate the influence of user pressure and contextual information, we propose SPAE (Suppressing Pressure, Amplifying Evidence), a training-free framework that uses the model's own judgments to identify relevant tokens, suppressing user pressure and amplifying contextual information through token-level attention steering. On average across five backbones, SPAE reduces pressure following by 18.8 percentage points and increases joint-condition updating by 5.5 percentage points relative to the strongest baseline in the main comparison. In two-turn dialogue, it improves joint-condition updating by an average of 13.2 percentage points over the strongest prompting baseline. The source data and codes can be found at https://github.com/03Grant/sycophancy-and-stubbornness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑