arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

拒绝的几何学:为何事后安全训练脆弱而预训练期间的安全持久

The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists

Srikanth Malla, Chiho Choi, Joon Hee Choi

arXiv 2609.06934首次发表:更新:

发表机构

Samsung Semiconductor US(三星半导体美国分部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过几何分析揭示事后安全训练脆弱而预训练期间安全持久的原因,并提出全程安全共训练方法,在多个模型上实现高攻击鲁棒性。

AI 中文摘要

事后安全训练(RLHF、DPO)是对齐大型语言模型的主要方式,然而越狱(Zou等人,2023b)、微调攻击(Qi等人,2024)以及激活空间探针(Arditi等人,2024)不断恢复出它本应移除的行为。我们为这种脆弱性给出一个几何解释,并追溯安全在预训练期间何时能够扎根。我们将安全更新 Δ = W_safe - W_base 与模型能力的曲率(能力损失的实证Fisher信息)进行对比。事后安全始终落入一种抑制机制:Δ 几乎与能力方向正交,其子空间内的微小部分集中在少数高曲率方向上。该更新薄而尖锐,是在完整能力之上铺设的拒绝门,而非对能力的擦除。一个核不动性引理解释了为何此类更新只能掩盖能力而不能移除它,因此少量良性微调即可恢复该能力:在Qwen-2.5-7B和Llama-3-8B Instruct上,100步良性微调使拒绝能力在保留能力上崩溃,这一特征在五个模型家族中复现。将该思路引入预训练,对OLMo-2-1B(OLMo等人,2025)的267个检查点扫描显示,安全所作用的基底在大约60亿到600亿预训练token之间以急剧转变的方式出现。随后我们建设性地利用这一思路:从零开始训练并全程持续进行安全共训练的模型,在各规模下达到87%至98%的拒绝率,攻击后水平保持在84%至91%,侵蚀仅为2至14个百分点,而事后安装的侵蚀为35至38个百分点,同时能力与仅语言模型基线相当或更优,且从4.1亿到69亿参数规模均保持这一表现,而计算量匹配的分段式调度则无法建立持久的拒绝。安全信号在预训练过程中的持续性,而非其时机,才是获得攻击鲁棒性的关键。

英文摘要

Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023), fine-tuning attacks (Qi et al., 2024), and activation-space edits (Arditi et al., 2024) keep recovering the behaviors it was meant to remove. We give this fragility one geometric explanation and follow it into pretraining. We measure the safety update $Δ= W_{safe} - W_{base}$ against the curvature of the model's capabilities (the empirical Fisher of a capability loss). Across five model families, post-hoc safety lands in a suppression regime: $Δ$ is nearly orthogonal to the capability directions, and its small in-subspace part concentrates on a few high-curvature ones. The update is thin but sharp, a refusal gate laid over intact capabilities rather than erasure of them. A kernel-immobility lemma explains why such an update can only mask a capability, not remove it, so a little benign fine-tuning restores it: 100 benign examples cut the AdvBench refusal of Qwen-2.5-7B-Instruct and Llama-3-8B-Instruct by 35 to 38 pp. Following the account into pretraining, a pretraining-checkpoint sweep of OLMo-2-1B (Team OLMo et al., 2024) shows the features that refusal attaches to emerging in a sharp transition between 1B and 63B pretraining tokens. We then use the account constructively: models trained from scratch with safety co-training spread continuously across pretraining reach 87 to 98% AdvBench refusal that the same attack erodes by only 2 to 14 pp at every scale from 410M to 6.9B, against 35 to 38 pp for post-hoc installs, at a small cost on short-answer capability probes; a windowed schedule of equal total safety weight installs no refusal. Persistence of the safety signal across pretraining, not its timing, is what buys attack robustness.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑