arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GAPS:用于条件激活引导的维度级门控

GAPS: Dimension-Level Gates for Conditional Activation Steering

Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad, A. B. Siddique

arXiv 2609.01878首次发表:更新:

发表机构

University of Kentucky; Dalhousie University(肯塔基大学; 达尔豪斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出GAPS方法,通过维度级门控补充现有条件激活引导方法,在毒性缓解与概念移除任务上,提升了语言模型的行为-能力权衡表现。

AI 中文摘要

激活引导通过在生成过程中向隐藏状态添加引导向量,来抑制语言模型中不合需要的行为。近期的条件方法如CAST和DSAS,通过决定何时进行干预,改善了行为与能力的权衡,但它们一旦激活,就会将完整的密集向量应用于所有隐藏维度,无论神经元是否携带概念信息,或是否已处于期望的状态。我们引入维度级条件作为补充的选择性轴,还决定要对哪些神经元进行干预。我们的方法GAPS(基于后验与可分性的门控激活引导)结合了两种无需训练的门控:静态可分性门控,它通过AUROC将引导限制在具有统计可靠概念信息的神经元上;以及动态后验门控,它仅在神经元当前激活能被高斯模型下的不合需要概念更好解释时,才对该神经元进行引导。这些门控每个令牌增加O(D)的开销,且可插入现有条件方法中。在使用Gemma-3(4B)和Qwen-3(1.7B)进行的毒性缓解(RealToxicityPrompts)和概念移除(OneSeC)任务上,GAPS始终与令牌级对应方法的帕累托前沿相匹配或有所提升;在固定能力预算下,DSAS+GAPS将Gemma-3的毒性率从6.52%降至0.48%,而仅使用DSAS时为3.52%。消融实验表明,大部分增益来自后验门控。

英文摘要

Activation steering suppresses undesired behaviors in language models by adding a steering vector to the hidden state during generation. Recent conditional methods such as CAST and DSAS improve the behavior-capability trade-off by deciding when to intervene, but once active, they apply the full dense vector to all hidden dimensions, regardless of whether a neuron carries concept information or already lies in the desired regime. We introduce dimension-level conditioning as a complementary axis of selectivity that also decides which neurons to intervene on. Our method, GAPS (Gated Activation steering via Posterior and Separability), combines two training-free gates: a static separability gate that restricts steering to neurons with statistically reliable concept information (via AUROC), and a dynamic posterior gate that steers a neuron only when its current activation is better explained by the undesired concept under a Gaussian model. The gates add O(D) overhead per token, and they plug into existing conditional methods. On toxicity mitigation (RealToxicityPrompts) and concept removal (OneSeC) with Gemma-3 (4B) and Qwen-3 (1.7B), GAPS consistently matches or improves the Pareto front of its token-level counterparts; under a fixed capability budget, DSAS+GAPS reduces Gemma-3's toxicity rate from 6.52% to 0.48%, versus 3.52% for DSAS alone. Ablations attribute most of the gain to the posterior gate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑