arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

注意力流形:通过编辑学习的B样条曲面来引导或阻断语言模型

Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces

Naveen Mysore

arXiv 2610.00257首次发表:更新:

发表机构

University of California, Santa Barbara; Dyssonance AI(加州大学圣塔芭芭拉分校; Dyssonance AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本工作提出注意力流形,即学习得到的B样条曲面,通过调制值维度来增强或阻断Transformer注意力,在降低困惑度的同时实现可编辑的模型引导与安全控制。

AI 中文摘要

在标准Transformer注意力机制中,源标记向每个接收者发送相同的值向量。查询决定“关注多少”,但不决定“提取什么”。本工作引入**注意力流形**:学习得到的二维B样条曲面$S_d(q_d, k_d)$,基于查询-键交互调制每个值维度。每个曲面是初始化为零的张量积三次B样条,从而保留预训练行为。将注意力流形应用于LLaMA 3.2-1B-Instruct和3B-Instruct,在0.3%参数开销下,将WikiText-2验证困惑度降低了2至2.5个点。在112个多样化提示中,曲面改变了69%(1B)至83%(3B)案例的贪心解码输出,对模糊和多义词输入的影响最强(94%至100%的变化率)。曲面提升了输出质量:纠正事实错误(“CAP定理有三个主要组成部分”变为“不可能同时保证全部三个”),提高精确性(“不可能知道某些属性”变为“不可能同时知道位置和动量”),并增加具体性(一个通用引用变为一个归因于圣奥古斯丁的引文,在两个规模上一致)。学习到的曲面也可机械编辑:反转某一层的系数会改变9/10提示的贪心输出(KL约0.010),为模型引导提供几何机制。将曲面系数设为-1会创建“注意力墙”,阻断特定维度的值流。在初步实验中,一层范围的墙将爆炸装置提示从具体指令重定向到一般教育内容,暗示了一条通往面向安全的流形塑造的路径。

英文摘要

In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \emph{how much} to attend but not \emph{what} to extract. This work introduces \textbf{attention manifolds}: learned 2D B-spline surfaces $S_d(q_d, k_d)$ that modulate each value dimension based on the query-key interaction. Each surface is a tensor-product cubic B-spline initialized to zero, preserving pretrained behavior. Applied to LLaMA 3.2-1B-Instruct and 3B-Instruct, attention manifolds reduce WikiText-2 validation perplexity by 2--2.5 points with 0.3\% parameter overhead. Across 112 diverse prompts, surfaces change greedy-decoded output for 69\% (1B) to 83\% (3B) of cases, with the strongest effects on ambiguous and polysemous inputs (94--100\% change rate). The surfaces improve output quality: correcting factual errors (\emph{``the CAP theorem has three main components''} $\to$ \emph{``it is impossible to guarantee all three''}), increasing precision (\emph{``impossible to know certain properties''} $\to$ \emph{``impossible to know both position and momentum''}), and adding specificity (a generic quote $\to$ an attributed Saint Augustine citation, consistently at both scales). The learned surfaces are also mechanically editable: inverting a layer's coefficients changes greedy output for 9/10 prompts (KL~0.010), providing a geometric mechanism for model steering. Setting surface coefficients to $-1$ creates ``attention walls'' that block value flow through specific dimensions. In a preliminary experiment, a layer-wide wall redirects an explosive-device prompt from specific instructions to general educational content, suggesting a path toward safety-oriented manifold shaping.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑