发表机构
POSTECH(浦项科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对语言模型的错误弃权(不执行)问题,将安全调优响应分解为模板化弃权声明与理由,发现仅基于理由训练可减少错误弃权且维持安全性能,为构建协调有用性与安全性的对齐智能体提供了方向。
AI 中文摘要
在对齐大型语言模型时,在有用性与安全性之间取得平衡仍是一项根本挑战。为实现这一平衡,模型应拒绝有害查询(例如“我如何射击他人?”),同时对良性输入保持响应,即使是那些表面上类似有害查询的良性输入(例如“我在哪里可以拍摄好照片?”)。然而,模型往往难以区分真正的有害查询与包含表面风险语言的良性查询,从而导致错误弃权(不执行)。在本文中,我们通过将安全调优数据集中的响应分解为两个不同组成部分来解决该问题:(i)模板化弃权(不执行)声明;(ii)解释弃权(不执行)的理由。我们的实验与分析表明,弃权(不执行)声明会诱导模型依赖表面线索,从而阻碍对有害与良性查询的准确区分。相比之下,仅基于理由进行训练可减少错误弃权(不执行),同时保持相当的安全性能。仅采用理由训练的优势在我们的上下文学习(ICL)配置中也有所体现,并且与所评估的推理时缓解方法兼容。这些结果强调了精心策划的细粒度安全监督数据集的必要性,并为构建能更好协调有用性与安全性的对齐智能体指明了方向。
英文摘要
Striking a balance between helpfulness and safety remains a fundamental challenge in aligning large language models. To achieve this balance, models should refuse harmful queries (e.g., "How do I shoot someone?") while remaining responsive to benign inputs, even those superficially resembling harmful queries (e.g., "Where can I shoot a good photo?"). However, models often struggle to distinguish genuinely harmful queries from benign queries that contain superficially risky language, resulting in false refusals. In this paper, we address the issue by decomposing a response in the safety-tuning dataset into two distinct components: (i) a boilerplate refusal statement and (ii) a rationale explaining the refusal. Our experiments and analyses show that refusal statements impede accurate discrimination between harmful and benign queries by inducing reliance on superficial cues. In contrast, training solely on rationales reduces false refusals while maintaining a comparable level of safety performance. Rationale-Only benefits also appear in our ICL configuration and remain compatible with the evaluated inference-time mitigation methods. The results emphasize the necessity of precisely curated, fine-grained safety supervision datasets and outline directions for constructing aligned agents that better reconcile helpfulness with safety.
CommentsEMNLP 2026 Main Conference (38 pages); Code available at https://github.com/mz-kim/RwR