arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Transformer 拒绝机制中的组件与维度稀疏性

Component and Dimension Sparsity in Transformer Refusal Mechanisms

Vincent Siu, Glenn Grant-Richards, Vlad Pavlovich, Yizhou Sun, Dawn Song, Chenguang Wang

arXiv 2610.06903首次发表:更新:

发表机构

UC Santa Cruz; UCLA; UC Berkeley(加州大学圣克鲁兹分校; 加州大学洛杉矶分校; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分解四个开源模型的拒绝引导,发现拒绝机制在组件和维度两个层面呈稀疏性,由可识别的结构化机制组装而成,而非分散编码。

AI 中文摘要

激活引导通过干预内部激活来操纵大型语言模型的行为,但这些干预的机制基础仍知之甚少。我们将拒绝引导分解为四个开源模型上的组件级干预,识别出注意力组件和 MLP 组件的稀疏子集,这些子集的引导足以复现完整的行为效果。我们发现拒绝方向集中在稀疏的组件机制中,该机制包含上游组件的 28% 至 48%,保留了引导有效性的 88% 至 101%。在这些机制内部,有效的引导进一步集中在残差流维度的大约 50% 中,保留了组件机制基线的 85% 至 98%,这与特权基结构一致。因此,稀疏性在两个层面上运作:哪些组件被引导,以及这些组件内的哪些维度携带信号。这些发现共同表明,拒绝并非分散地编码在 Transformer 中,而是由结构化的、可识别的机制组装而成,为理解拒绝行为如何被表示和引导的机制提供了基础。为促进可复现性,我们在此 https URL 发布了所有代码和原始实验结果。

英文摘要

Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in sparse component mechanisms comprising 28--48\% of upstream components, retaining 88--101\% of steering effectiveness. Within these mechanisms, effective steering further concentrates in approximately 50\% of residual stream dimensions, retaining 85--98\% of the component-mechanism baseline, consistent with a privileged basis structure. Sparsity thus operates at two levels: which components are steered, and which dimensions within those components carry the signal. Together these findings show that refusal is not diffusely encoded across a transformer but assembled by a structured, identifiable mechanism, providing a foundation for mechanistic understanding of how refusal behaviors are represented and steered. To facilitate reproducibility, we release all code and raw experimental results in https://github.com/wang-research-lab/Refusal_Mechanisms.

CommentsAccepted to COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑