发表机构
The University of Hong Kong; Shenzhen Loop Area Institute; Shanghai Artificial Intelligence Laboratory(香港大学; 深圳河套学院; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过补偿性特征注入减少模型谄媚行为,发现其不总能增强直接拒绝,但在用户压力下可部分恢复拒绝能力,为AI安全训练提供新见解。
AI 中文摘要
可靠地拒绝有害请求对于语言模型的安全部署至关重要。由于过度取悦用户的倾向可能削弱现有的拒绝能力,减少谄媚为在安全训练所覆盖的有害场景之外增强拒绝能力提供了一条潜在途径。我们使用补偿性特征注入(CFI)这一训练技术来研究这种可能性,该技术通过在训练期间提供目标概念的相关激活来限制其习得。在三个Qwen3.5基础模型上,我们使用稀疏自编码器(SAEs)从配对的谄媚响应和独立响应中识别排名最高的谄媚特征,然后通过推理引导验证其行为影响。随后,我们在针对谄媚目标的监督微调期间注入所选特征。正向注入在移除后减少了习得的谄媚行为(在35B-A3B中相对于普通微调降低了62.0%),而适度负向注入则增加了谄媚行为。出乎意料的是,这些谄媚行为的减少并未一致地改善对有害请求的直接拒绝,这促使我们在用户压力下对相同的有害意图进行更窄范围的评估。在这种设置下,针对谄媚响应的普通微调显著削弱了拒绝能力,而使用正向注入训练的选定检查点恢复了部分损失,在35B-A3B中恢复了约95%。这些发现表明,持续的谄媚减少并不能保证更强的直接拒绝,同时识别出在用户压力下的恢复是训练干预的一个独特、有条件的益处。
英文摘要
Reliable refusal of harmful requests is essential to the safe deployment of language models. Because excessive eagerness to please users may undermine existing refusal capabilities, reducing sycophancy offers a potential route to stronger refusal beyond the harmful scenarios covered by safety training. We investigate this possibility using compensatory feature injection (CFI), a training technique designed to limit the acquisition of a target concept by supplying its associated activation during learning. Across three Qwen3.5 base models, we use sparse autoencoders (SAEs) to identify the top-ranked sycophancy feature from paired sycophantic and independent responses, then validate its behavioral influence through inference steering. We subsequently inject the selected feature during supervised fine-tuning on sycophantic targets. Positive injection reduces learned sycophancy after removal (by 62.0% relative to ordinary fine-tuning in 35B-A3B), whereas modest negative injection increases it. Unexpectedly, these reductions in sycophancy do not consistently improve direct refusal of harmful requests, motivating a narrower evaluation of the same harmful intents under user pressure. In this setting, ordinary fine-tuning on sycophantic responses substantially weakens refusal, while selected checkpoints trained with positive injection recover part of the loss, including approximately 95% in 35B-A3B. These findings show that persistent sycophancy reduction does not guarantee stronger direct refusal, while identifying recovery under user pressure as a distinct, conditional benefit of training intervention.
Comments20 pages