arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11583cs.AI

安全对齐的定位:MLP层与网络中间块编码大语言模型的拒绝行为

Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models

Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过移植不同粒度的模型权重实验,发现大语言模型的安全拒绝行为主要编码在MLP层,且集中于网络中间块,其构成具有非加性,存在基准依赖的精度-覆盖权衡。

中文摘要 AI 辅助

大语言模型的安全对齐常被视为整个网络的分布式属性,但其实际的脆弱性表明,拒绝行为可能集中在一组更小的参数中。本研究通过将对齐模型的权重以多种粒度移植到匹配的未对齐基础模型中,探究安全对齐的拒绝行为编码位置。使用两组开源权重模型对和四个安全基准,我们开展实验比较替换注意力权重、MLP权重、连续层区域及MLP块的效果。在两个模型族中,拒绝行为的迁移由MLP权重主导:替换MLP参数相比替换注意力参数,能恢复多得多的恶意提示拒绝行为,在各基准上增益至少达2.7倍。在MLP栈内,与拒绝相关的参数呈现一致的网络中间层集中特征,在所有6次模型-数据集对的贪心搜索中,层8-11组成的块均被优先选中。结果还显示,安全相关组件的构成具有非加性:在6条贪心轨迹中的5条里,添加更多对齐块会降低拒绝性能,而选择性块子集在恶意拒绝、良性过度拒绝或两者上的表现优于完整MLP移植。最后,迁移到OR-Bench的贪心顺序随用于推导它的源基准而变化,表明存在依赖基准的精度-覆盖权衡。这些结果表明,当前大语言模型的安全对齐既是局部化的,又具有交互敏感性,为理解对齐脆弱性及靶向安全干预的潜在途径提供了见解。

英文摘要

Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters. This work addresses where safety-aligned refusal is encoded by transplanting weights from aligned models into matched unaligned base models at multiple levels of granularity. Using two open-weight model pairs and four safety benchmarks, we conducted experiments to compare the effects of replacing attention weights, MLP weights, contiguous layer regions, and MLP blocks. Across both model families, refusal transfer is dominated by MLP weights: replacing MLP parameters recovers substantially more malicious-prompt refusal than replacing attention parameters, with gains of at least 2.7 times more across benchmarks. Within the MLP stack, refusal-relevant parameters exhibit a consistent mid-network concentration, as the block spanning layers 8-11 is selected first in all six greedy searches over model-dataset pairs. The results also show that the composition of safety-relevant components is non-additive: in five of six greedy trajectories, adding more aligned blocks can reduce refusal performance, and selective block subsets can outperform full MLP transplantation on malicious refusal, benign over-refusal, or both. Finally, greedy orders transferred to OR-Bench vary with the source benchmark used to derive them, indicating a benchmark-dependent precision-coverage trade-off. These results suggest that safety alignment in current LLMs is both localized and interaction-sensitive, offering insight into alignment brittleness and potential avenues for targeted safety interventions.

发表机构

  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

↑