arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36254cs.AI

缓解大型推理模型中的欺骗性安全对齐

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

Xiangyu Zhou, Saleh Zare Zade, Rafi Ibn Sultan, Alexander Kotov, Dongxiao Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对大型推理模型因RL仅奖励最终答案导致推理与答案安全信号不一致的欺骗性安全对齐问题,提出DSAR指标量化及SARA强化学习方法,在保持有用性下显著缓解该现象。

中文摘要 AI 辅助

大型推理模型(LRMs)通常使用强化学习(RL)进行训练,以改进其在生成最终答案之前的思维链(CoT)推理过程。然而,RL奖励通常基于最终答案进行分配,对中间推理过程几乎没有直接监督。这可能导致欺骗性安全对齐,即推理轨迹与最终答案传达不一致的安全信号。为系统研究该现象,我们引入DSAR(欺骗性安全对齐率),一种联合评估推理轨迹与最终答案以量化其安全不一致性的指标。在多个LRM和基准测试中,我们发现欺骗性安全对齐在标准提示条件下普遍存在,并在预填充攻击下显著加剧。我们进一步提供隐藏表示分析,表明模型在最终答案阶段比在中间推理阶段表现出更强的安全区分能力。为缩小这一差距,我们提出SARA(安全感知推理对齐),一种基于RL的方法,同时奖励安全感知推理和安全最终答案,鼓励早期有害意图识别并强制推理-答案一致性。实验表明,SARA在标准和对抗性设置下均显著缓解欺骗性安全对齐,同时保持有用性和实用性。代码可在该https URL获取。

英文摘要

Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility. Code is available at https://github.com/xzhou98/SARA.

发表机构

  • Wayne State University(韦恩州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑