AUDITPLAN:先承诺,再回答,实现可审计的安全对齐
AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment
浏览论文内容
中文总结 AI 辅助
针对安全微调难以区分稳健拒答与不良捷径的问题,提出AUDITPLAN单模型先计划后回答方法,通过FAITHGATE奖励门控强化计划-答案耦合,显著提升安全对齐的忠实性、稳健性与可审计性。
中文摘要 AI 辅助
安全微调流程仅评判最终答案,这使得难以区分稳健的拒答与两种不良捷径:对良性请求的一概拒答,以及看似合理但实际并未约束答案的不忠实安全理由。我们提出AUDITPLAN,一种单模型“先计划后回答”的方法,模型首先输出一个紧凑的结构化安全计划,然后基于该计划生成答案。该计划记录威胁标签、预期行动和明确约束,支持机器可检查的审计,同时在部署时对用户隐藏。我们通过监督微调,随后使用FAITHGATE进行强化学习来训练此行为,FAITHGATE是一种奖励门控目标,仅在安全计划正确时才授予答案奖励。这抑制了看似安全但不忠实的行为,并促进计划与答案之间更紧密的耦合。在Qwen骨干模型上,AUDITPLAN同时提升了稳健性和可审计性:在Qwen2.5-3B-Instruct上,FAITHGATE将ASR从24.0%降至11.6%,LSR从1.0%降至0.36%,过度拒答从11.0%降至2.0%,优于仅答案的强化学习、自由形式解释和加权求和的结构化奖励。Qwen2.5-1.5B-Instruct上也呈现类似趋势。在Qwen-3-4B-Instruct和Qwen2.5-7B-Instruct上的更大模型确认运行保持了相同趋势,表明显式内部承诺可使安全对齐更忠实、更稳健且更可审计。
英文摘要
Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remaining hidden from users at deployment. We train this behavior with supervised fine-tuning followed by reinforcement learning with FAITHGATE, a reward-gating objective that grants answer reward only when the safety plan is correct. This discourages safe-looking but unfaithful behavior and promotes tighter plan-answer coupling. Across Qwen backbones, AUDITPLAN improves both robustness and auditability: on Qwen2.5-3B-Instruct, FAITHGATE reduces ASR from 24.0% to 11.6%, LSR from 1.0% to 0.36%, and over-refusal from 11.0% to 2.0%, outperforming answer-only RL, free-form explanation, and weighted-sum structured rewards. Similar trends hold for Qwen2.5-1.5B-Instruct. Larger-model confirmation runs on Qwen-3-4B-Instruct and Qwen2.5-7B-Instruct preserve the same trend suggesting that explicit internal commitments can make safety alignment more faithful, robust, and auditable.
发表机构
- Indian Institute of Technology Patna(印度理工学院巴特那分校)
机构由 AI 辅助整理,请以论文原文为准。