发表机构
Beihang University; University of Science and Technology Beijing; National University of Singapore; Lanzhou University; Jinan University(北京航空航天大学; 北京科技大学; 新加坡国立大学; 兰州大学; 暨南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对情绪表达导致LLM对良性请求过度拒绝的问题,提出EmoRSS激活转向方法,通过构建拒绝子空间并施加反向干预,在保持有害请求拒绝的同时缓解过度拒绝,提升权衡与通用性能。
AI 中文摘要
情绪表达可以影响大型语言模型(LLM)的安全决策,为改进安全对齐提供了潜在途径。现有研究主要关注情绪表达如何在有害请求下促进攻击,而忽略了其对良性请求的影响。我们发现,情绪表达也会系统性地增加对良性请求的拒绝倾向,导致不必要的过度拒绝。基于这一观察,我们提出了情绪引导的拒绝子空间转向(EmoRSS),一种激活转向方法,在保留对有害请求的拒绝行为的同时,缓解情绪引发的过度拒绝。具体而言,我们首先使用逐层线性探针识别对拒绝敏感的网络层,并从与探针方向对齐的稀疏自编码器(SAE)特征中构建拒绝子空间。接下来,我们使用具有相同查询的成对常规和情绪化请求来估计定义拒绝子空间的特征中的平均激活偏移。最后,我们将此偏移解码为激活干预向量,并在推理期间沿反向拒绝方向应用它,而不更新骨干参数。在两个LLM上的实验表明,当请求包含情绪表达时,与先前的过度拒绝缓解基线相比,我们的方法在拒绝有害请求和回答良性请求之间实现了更有利的权衡,同时更好地保留了通用任务性能。
英文摘要
Emotional expression can influence the safety decisions of large language models (LLMs), offering a potential avenue for improving safety alignment. Existing studies have mainly focused on how emotional expressions facilitate attacks under harmful requests, while overlooking their effects on benign requests. We find that emotional expression can also systematically increase refusal tendencies on benign requests, leading to unnecessary over-refusal. Based on this observation, we propose emotion-guided refusal subspace steering (EmoRSS), an activation-steering method that mitigates emotion-induced over-refusal while preserving refusal behaviour on harmful requests. Specifically, we first identify a refusal-sensitive layer using layer-wise linear probes and construct a refusal subspace from sparse autoencoder (SAE) features aligned with the probe direction. Next, we use paired regular and emotional requests with the same queries to estimate the mean activation shift in the features defining the refusal subspace. Finally, we decode this shift into an activation intervention vector and apply it in the reverse refusal direction during inference, without updating the backbone parameters. Experiments on two LLMs show that, when requests contain emotional expressions, our method achieves a more favourable trade-off between refusing harmful requests and answering benign ones than prior over-refusal mitigation baselines, while better preserving general task performance.