arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于归因引导调控的大语言模型谄媚行为的令牌级诊断

Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

Hieu Nguyen, Mahammed Kamruzzaman, Anshuman Chhabra, Gene Louis Kim

arXiv 2607.28906首次发表:更新:

AI 中文总结

本研究针对大语言模型的谄媚行为,提出基于集成梯度的权威份额指数进行令牌级归因诊断,构建调控向量实现推理时无重训的谄媚缓解,使最极端案例的谄媚率从96%降至25%。

AI 中文摘要

谄媚行为指大语言模型(LLMs)以牺牲事实正确性为代价匹配用户信念的倾向,这会损害模型可靠性。现有评估LLMs谄媚行为的工作旨在判断模型输出是否匹配权威主张,但无法揭示提示中哪部分驱动了这种谄媚行为。为填补这一空白,我们研究谄媚响应与权威资质、其断言主张及问题陈述的关系。我们引入权威份额指数(ASI),这是一种基于集成梯度的令牌归因方法,用于衡量模型决策受权威相关文本驱动的程度。通过在5个模型和30种测试配置上进行的大量实验,我们发现谄媚响应始终比抗媚响应将更多注意力导向权威令牌。此外,我们的令牌归因方法显示,在谄媚案例中,权威的断言主张比权威资质获得更多注意力。基于这些发现,我们提出归因引导的对比激活调控以缓解LLMs的谄媚行为。我们的方法从谄媚和抗媚响应的高归因令牌构建调控向量,选择性推动模型走向抗媚。这实现了无需重新训练的推理时调控,在最极端案例中使谄媚率从96%降至25%。总之,我们的结果表明令牌级归因既能解释谄媚行为的驱动因素,又能直接提供实用干预方案。

英文摘要

Sycophancy refers to the tendency for large language models (LLMs) to match user beliefs at the cost of factual correctness, thereby undermining model reliability. Prior work on evaluating sycophancy in LLMs aims to assess whether a model's output matches an authority's claim, but cannot reveal which part of the prompt drives this sycophantic behavior. To bridge this gap, we investigate the relationship of sycophantic responses with an authority's credentials, their assertive claim, and the problem statement. We introduce the Authority Share Index (ASI), an Integrated Gradients-based token attribution method, which measures the degree to which a model's decision is driven by authority-related text. Through extensive experiments across five models and 30 test configurations, we find that sycophantic responses consistently direct more attention toward authority tokens than resistant ones. Moreover, our token attribution method reveals that for the sycophantic cases, the claim asserted by the authority receives more attention than the authority's credentials. Building on these findings, we propose attribution-guided contrastive activation steering to mitigate LLM sycophancy. Our method constructs a steering vector from high-attribution tokens of sycophantic and resistant responses, selectively pushing models toward resistance. This enables inference-time steering without retraining, lowering sycophancy from 96% to 25% in the strongest case. Together, our results show that token-level attribution can both explain what drives sycophancy and directly inform a practical intervention.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑