arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越浅层对齐:后训练方法如何决定拒绝回路与操控鲁棒性

Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness

Hoang Cuong Nguyen, Mark Dras, Usman Naseem

arXiv 2609.03887首次发表:更新:

发表机构

Macquarie University(麦考瑞大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比三种后训练方法在三款模型上的拒绝机制,发现训练方法与架构分别影响拒绝计算与操控性,指出无方法同时满足安全对齐的三项理想特性,提醒勿将后训练视为可靠防御。

AI 中文摘要

训练语言模型以拒绝有害请求的方法,是如何决定模型内部拒绝机制的运作方式的?我们在三种架构不同的模型(Llama-3.1-8B、Gemma-2-9B、Qwen3-8B)上,对比了三种后训练方法:监督微调、推理增强微调(基于为安全决策提供依据的推理链进行训练)、偏好优化(ORPO)。我们发现,训练方法而非仅数据,会重塑内部拒绝的计算方式:推理增强训练会产生一种独特的拒绝计算,在所有三种模型中均可见;而架构会独立影响内部结构及拒绝可被可靠操控的程度。最重要的是,我们研究的所有方法均未同时实现安全对齐所需的三个理想特性:拒绝不集中于少数脆弱组件、安全提升不损耗通用能力、安全行为可通过小型针对性编辑修正。我们提醒,勿将当前后训练方法视为已解决的可靠防御,尤其在安全关键的应用场景中。代码与模型可在此httpsURL获取。

英文摘要

How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervised fine-tuning, reasoning-augmented fine-tuning (training on reasoning chains that justify a safety decision), and preference optimization (ORPO) - across three architecturally distinct models (Llama-3.1-8B, Gemma-2-9B, Qwen3-8B). We find that training method, not just data, reshapes how refusal is computed internally: reasoning-augmented training consistently produces a distinct kind of refusal computation, visible across all three models, while architecture independently shapes internal structure and how reliably refusal can be steered. Most importantly, no method we study achieves all three properties we would want from safe alignment at once: refusal that isn't concentrated in a few fragile components, safety gains that don't cost general capability, and safety behavior correctable through small, targeted edits. We caution against treating current post-training methods as a solved, reliable defense, especially for security-critical use. Code and models are available in https://github.com/hoangcuongnguyen2001/Beyond-Shallow-Alignment.

Comments27 pages, accepted at EMNLP 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑