发表机构
Nara Institute of Science and Technology(奈良科学技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出跨编码导向评估方法,发现激活导向主要追踪提取索引而非语义标签,且不同数据集、模型及评估类型的表现存在差异,单一编码下的导向增益无法确定干预的控制对象。
AI 中文摘要
激活导向(Activation steering)通常是在用于构建方向的答案编码下进行评估的,所报告的增益可能反映了预期判断或构建过程中所见答案标识符的兼容性。我们提出跨编码导向评估(Cross-Encoding Steering Evaluation),该方法在冻结干预的同时,对保留项的答案进行重新编码。在NormBank数据集上,当A/B/C标识符被重新分配后,对比激活加法(CAA)在提取索引上诱导的目标与源得分变化,大于新映射下语义标签上的变化,我们将此称为“提取索引跟随”。改变标识符词汇(A/B/C、X/Y/Z或1/2/3)和行顺序,显示该效应追踪的是提取索引而非行位置。在匹配各层的方向范数后,提取索引跟随主要出现在更深层;包含方向平方范数15.4%的低秩输出敏感成分保留了96.3%的该效应。一种推理时干预(ITI)式方法在NormBank的三个模型上也表现出偏向提取索引而非语义标签跟随的情况。总体而言,MNLI偏向提取索引跟随,而Social Chemistry 101(SC101)偏向语义标签跟随;多项选择与开放式评估可能产生不同的行为结论。因此,单一答案编码下的导向增益本身无法识别该干预所控制的内容。
英文摘要
Activation steering is often evaluated under the answer encoding used to construct the direction. A reported gain may reflect the intended judgment or compatibility with answer identifiers seen during construction. We introduce Cross-Encoding Steering Evaluation, which freezes an intervention while re-encoding answers to the same held-out items. On NormBank, after A/B/C identifiers are reassigned, contrastive activation addition (CAA) induces larger target-versus-source score changes for the extraction indices than for the semantic labels under the new mapping. We call this extraction-index following. Varying identifier vocabulary (A/B/C, X/Y/Z, or 1/2/3) and row order shows that the effect tracks extraction index rather than row position. After matching direction norms across layers, extraction-index following emerges mainly at later depths. A low-rank output-sensitive component containing 15.4% of the direction's squared norm retains 96.3% of this effect. An Inference-Time Intervention (ITI)-style method also favors extraction-index over semantic-label following on NormBank in three models. In aggregate, MNLI favors extraction-index following, whereas Social Chemistry 101 (SC101) favors semantic-label following. Multiple-choice and open-ended evaluations can yield different behavioral conclusions. Thus, a steering gain under one answer encoding does not by itself identify what the intervention controls.