IDEEA:基于激活簇匹配的无训练输入依赖引导
IDEEA: training-free Input-Dependent stEEring via Activation cluster matching
浏览论文内容
中文总结 AI 辅助
本研究针对现有无训练引导方法输入无关的局限,提出IDEEA框架,通过聚类注意力头的激活并匹配输入激活选择对应方向,在保留输入表示的同时对齐LLMs,使TruthfulQA的truth×info指标平均提升9.9%。
中文摘要 AI 辅助
引导(Steering)通过在推理时向选定的激活中注入偏差来对齐大语言模型(LLMs),相较于监督微调或强化学习等权重更新方法,成本低得多。然而,现有的大多数无训练引导方法都是输入无关的:仅拟合一个单一方向并在所有输入间共享。这存在根本局限,因为不同输入在激活空间中占据不同区域,针对同一目标概念会有不同的最优引导方向,就像固定损失对应的梯度会随输入变化一样。我们提出IDEEA(Input-Dependent stEEring via Activation cluster matching,基于激活簇匹配的输入依赖引导),一种用于输入依赖引导的无训练框架。IDEEA对每个注意力头的正负激活支持进行聚类,通过求解最优匹配问题构建一组与目标概念相关的簇条件方向;在推理时,从该方向池中选择与输入自身激活最匹配的方向进行引导。IDEEA在将模型对齐至目标概念的同时,保留输入的原始表示,这证明编码某一概念的激活在表示空间中占据多个不同子区域而非单一区域。在TruthfulQA数据集上,IDEEA的truth×info指标较最优的输入无关基线平均提升9.9%,最高提升23.5%。
英文摘要
Steering aligns large language models (LLMs) by injecting a bias into selected activations at inference time, offering a far cheaper alternative to weight-update methods such as supervised fine-tuning or reinforcement learning. However, most existing training-free steering methods are input-independent: a single direction is fitted once and shared across all inputs. This is fundamentally limiting as different inputs occupy different regions of the activation space and admit different optimal steering directions toward the same target concept, much as the gradient with respect to a fixed loss varies from input to input. We close this gap with IDEEA (Input-Dependent stEEring via Activation cluster matching), a training-free framework for input-dependent steering. IDEEA clusters the positive and negative activation supports per attention head, and solves an optimal-matching problem to construct a set of cluster-conditional directions, all about the target concept. At inference time, it picks from this pool of directions and uses the one that best matches the input's own activation for steering. IDEEA aligns the model toward the target concept while preserving the input's original representation, evidence that activations encoding a concept occupy several distinct sub-regions of the representation space rather than a single one. IDEEA improves the truth $\times$ info rate in TruthfulQA by an average of 9.9% (up to 23.5%) over the best input-independent baseline.
发表机构
- University of Waterloo(滑铁卢大学)
- University of British Columbia(不列颠哥伦比亚大学)
- Vector Institute for AI(人工智能矢量研究所)
- University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。