发表机构
The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对LLMs的谄媚性问题,提出PCA引导的激活缩放框架,通过分解残差流激活实现单调双向控制,在三个模型和三个数据集上的表现优于基线方法。
AI 中文摘要
大型语言模型(LLMs)存在谄媚性,即无论事实准确性如何都倾向于同意用户观点,这可能强化误解,但完全消除谄媚性又可能导致对有效观点的过度修正。因此,有效的控制必须既能降低又能提升谄媚性,且效果可预测且渐进。然而,现有方法无法确保在不同模型和数据集上,控制强度与行为结果之间存在双向且单调的关系。我们提出PCA引导的激活缩放(PAS),这是一种激活控制框架,它将残差流激活分解为经PCA识别的谄媚性-诚实性子空间和正交残差,随后应用不同的缩放指数以实现单调、双向的控制。在三个LLMs和三个数据集上,PAS实现了强单调性(斯皮尔曼相关系数ρ=+0.92),且每个方向的平均偏移为15.4%,而基线方法仅为8.7%。消融研究证实,分解、非对称指数和层选择对于维持单调控制均至关重要。数据和代码可在该https网址获取。
英文摘要
Large language models (LLMs) exhibit sycophancy, a tendency to agree with user beliefs regardless of factual accuracy. This can reinforce misconceptions, but eliminating it entirely risks over-correction against valid opinions. Effective control must therefore both reduce and increase sycophancy with predictable and gradual effect. Yet, existing methods fail to ensure a bidirectional and monotonic relationship between steering strength and behavioral outcome across models and datasets. We introduce PCA-guided Activation Scaling (PAS), an activation steering framework that decomposes residual stream activations into a PCA-identified sycophancy-honesty subspace and an orthogonal residual, then applies distinct scaling exponents to achieve monotonic, bidirectional control. Across three LLMs and three datasets, PAS achieves strong monotonicity (Spearman $ρ$ = +0.92) and an average shift of 15.4% per direction, compared with 8.7% for the baselines. Ablation studies confirm that the decomposition, asymmetric exponents, and layer selection are each essential for maintaining monotonic control. The data and code are available at https://github.com/Bellafc/PCS.
Commentsaccepted by COLM2026