arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过影响力引导:用于激活引导的曲率感知数据加权

Steering by Influence: Curvature Aware Data Weighting for Activation Steering

James A. E. Dixon, Stephen J. Roberts, Francesco Quinzan

arXiv 2610.06383首次发表:更新:

发表机构

University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出基于影响力函数的曲率感知数据加权方法,用于激活引导,通过最优传输加权概念示例,在毒性抑制、概念归纳和真实性任务上优于基线并保持模型质量。

AI 中文摘要

推理时引导通过在激活空间中估计概念的表征并朝向该表征移动激活,提供了一种廉价且细粒度的语言模型输出控制方法。现有方法基于对比数据集上的激活平均值构建这些表征。这些平均值包含了不相关的概念和噪声,并且由少数标记主导,这意味着激活迁移编码的是标记级而非主题级概念。在本工作中,我们朝向最能以主题方式表达概念的示例进行引导,而不是朝向所有示例的期望。我们使用影响力函数来识别这些示例,该函数估计每个数据点对模型概念表征的贡献程度。与简单的模型激活相似性不同,影响力函数考虑了模型损失景观的曲率,从而能够捕捉超越表面标记级相似性的概念相关关系。随后,我们提出了影响力加权激活迁移,该方法使用最优传输将非概念文本的激活引导至概念文本的激活,并根据影响力分数对概念示例进行加权。我们在毒性抑制(Jigsaw)、基于对象的概念归纳(OneSec)和真实性归纳(TruthfulQA)上进行了评估,优于现有的激活迁移基线。我们使用困惑度和MMLU准确率跟踪引导后的能力,发现我们的方法在改善引导的同时基本保持了模型质量。我们进一步表明,影响力函数捕捉了基于激活的方法所遗漏的概念相关信息,两种方法对数据点的排序显著不同。综合这些结果,展示了曲率感知的影响力信息对于激活引导的价值。

英文摘要

Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages over contrastive datasets. These averages incorporate unrelated concepts and noise, and are dominated by a few tokens, meaning the activation transport encodes token-level rather than thematic concepts. In this work, we steer towards examples that most express a concept thematically, rather than towards an expectation over all. We identify these examples using influence functions, which estimate how much each data point contributes to a model's representation of a concept. Unlike simple model activation similarity, they incorporate the curvature of the model's loss landscape, allowing them to capture concept-relevant relationships beyond superficial token-level similarity. We then propose influence-weighted activation transport, which uses optimal transport to steer activations of non-concept text towards those of concept text, weighting concept examples by their influence scores. We evaluate on toxicity suppression (Jigsaw), object-based concept induction (OneSec) and truthfulness induction (TruthfulQA), outperforming existing activation-transport baselines. We track capability after steering using perplexity and MMLU accuracy, finding that our method improves steering while largely preserving model quality. We further show that influence functions capture concept-relevant information that activation-based methods miss with the two approaches ranking data points significantly differently. Together, these results demonstrate the value of curvature-aware influence information for activation steering.

CommentsCode: https://github.com/JDIXON-2/Concept_Activation_Transport

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑