arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.23260cs.CLcs.AIcs.LG

通过SAE构造的低秩子空间适应实现可解释的安全对齐

Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation

  • Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

Dianyun Wang, Qingsen Ma, Yuhu Shang, Zhifeng Lu, Zhenbo Xu, Lechen Ning, Huijia Wu, Zhaofeng He

更新

AI总结:

SAILS通过SAE构造低秩子空间实现安全对齐,提升安全性能并保持参数效率。

AI中文摘要:

安全对齐——训练大型语言模型(LLMs)以拒绝有害请求同时保持有帮助——对于负责任的部署至关重要。先前的工作确立了安全行为由低秩结构支配,建议参数高效微调(PEFT)应适合对齐。然而,低秩适应(LoRA)在安全基准上始终表现不佳,比全微调和强化学习差。我们归因于语义缠结:安全相关方向与无关概念交织在一起,由于多义性,阻碍了隐式子空间识别。为了解决这个问题,我们提出了SAILS(通过可解释低秩子空间实现安全对齐),它利用稀疏自动编码器(SAEs)将表示分解为单义特征,从SAE解码器方向构建可解释的安全子空间,并用它来初始化LoRA适配器。理论上,我们证明了基于SAE的识别在单义性假设下可以实现任意小的恢复误差,而直接识别则有不可减少的误差底限。经验上,SAILS在Gemma-2-9B上实现了高达99.6%的安全率——比全微调高7.4个百分点,与基于RLHF的模型相当——同时仅更新0.19%的参数并提供可解释性。

英文摘要:

Safety alignment -- training large language models (LLMs) to refuse harmful requests while remaining helpful -- is critical for responsible deployment. Prior work established that safety behaviors are governed by low-rank structures, suggesting parameter-efficient fine-tuning (PEFT) should be well-suited for alignment. However, Low-Rank Adaptation (LoRA) consistently underperforms full fine-tuning and reinforcement learning on safety benchmarks. We attribute this gap to semantic entanglement: safety-relevant directions are intertwined with unrelated concepts due to polysemanticity, impeding implicit subspace identification. To address this, we propose SAILS (Safety Alignment via Interpretable Low-rank Subspace), which leverages Sparse Autoencoders (SAEs) to disentangle representations into monosemantic features, constructs an interpretable safety subspace from SAE decoder directions, and uses it to initialize LoRA adapters. Theoretically, we prove that SAE-based identification achieves arbitrarily small recovery error under monosemanticity assumptions, while direct identification suffers an irreducible error floor. Empirically, SAILS achieves up to 99.6% safety rate on Gemma-2-9B -- exceeding full fine-tuning by 7.4 points and matching RLHF-based models -- while updating only 0.19% of parameters and providing interpretability.

↑