arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31122cs.CLcs.LG

LocUS:用于目标激活引导的头部选择与子空间投影

LocUS: Head Selection and Subspace Projection for Targeted Activation Steering

  • Université Côte d’Azur(蔚蓝海岸大学)
  • Inria(法国国家信息与自动化研究所)
  • Area Science Park(的里雅斯特科技园区)
  • Institute of Science Technology Austria(奥地利科学技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

Irene Tallini, Lorenzo Basile, Valentino Maiorca, Francesco Locatello, Alberto Cazzaniga

AI总结:

本文提出LocUS方法,通过选择注意力头并投影到非嵌入子空间,实现目标激活引导,在干预少量参数下达到或超越现有基线,并保持通用能力。

AI中文摘要:

激活引导是一种强大的无需训练的范式,用于在推理时控制大型语言模型。然而,标准方法从对比数据中估计每层的引导方向,并将其应用于该层的整个表示空间,这可能会将干预与对比数据中存在的非目标属性耦合,从而降低无关能力。为缓解这一问题,我们引入了LocUS(局部化非嵌入引导),该方法将激活引导基于模型自身的输出词汇子空间。通过在非嵌入矩阵中识别特定于属性的线性子空间,LocUS施加了几何约束,将引导变换限制在特定子空间内,同时将其应用局部化到稀疏的注意力头子集。在三个模型家族上针对毒性缓解、情感重定向和谄媚抑制的广泛评估表明,LocUS在干预少于6%的参数的同时,匹配或优于最先进的基线,并更好地保持通用能力。

英文摘要:

Activation steering is a powerful training-free paradigm for controlling large language models at inference time. However, standard approaches estimate a per-layer steering direction from contrastive data and apply it on the layer's entire representation space, which may couple the intervention to off-target properties present in the contrastive data and degrade unrelated capabilities. To mitigate this issue, we introduce LocUS (Localized Unembedding Steering), a method which grounds activation steering to the model's own output vocabulary subspace. By identifying a property-specific linear subspace within the unembedding matrix, LocUS enforces a geometric constraint that restricts the steering transformation to a specific subspace and at the same time localizes its application to a sparse subset of attention heads. Extensive evaluations across three model families on toxicity mitigation, sentiment redirection and sycophancy suppression show that LocUS matches or outperforms state-of-the-art baselines while intervening on under 6% of parameters and better preserving general capability.

↑