arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

区分大语言模型中谄媚行为的内部表征

Dissociating the Internal Representations of Sycophancy in LLMs

Anthony Baez, Sheer Karny, Pat Pataranutaporn

arXiv 2607.07003首次发表:更新:

AI 中文总结

研究大语言模型谄媚行为,将其表征分为事实和观点子类型,通过训练线性探针等方法评估不同模型对两子类型表征差异,为研究复杂模型行为表征结构提供了新框架。

AI 中文摘要

大语言模型(LLMs)经常表现出谄媚行为,即即便用户陈述错误也会认同。谄媚常被视为单一行为,但它能以多种不同方式和情境呈现,这引发了其多面性是否在内部机制中体现的问题。为填补这一空白,我们将谄媚表征分为事实和观点子类型,受可验证声明和主观信念差异驱动。我们训练线性探针并基于一种子类型的激活构建引导向量,评估其向另一子类型的转移,以衡量它们共享表征的程度。我们发现不同的大语言模型对这些子类型的表征不同,有的表征更统一,有的更不同且存在因果干扰。这种区分方法为研究复杂模型行为的表征结构提供了一个有前景的框架。

英文摘要

Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophancy can manifest in substantially distinct ways across contexts, raising the question of whether this heterogeneity is reflected in its internal mechanisms. To address this gap, we dissociate the representations of sycophancy into factual and opinion subtypes, motivated by prior evidence of heterogeneous truth representations in LLMs. We train linear probes and construct steering vectors on one subtype's activations and evaluate their transfer to the other, measuring the extent to which representations are shared and visualizing them via Linear Discriminant Analysis. We find that different LLMs represent these subtypes differently, with either more aligned or more distinct representations, and apply this insight to improve representational interventions for reducing sycophancy. Our dissociation method offers a general framework for studying the representational structure of complex model behaviors.

CommentsAccepted to Mechanistic Interpretability Workshop at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑