发表机构
State Key Laboratory of Multimedia Information Processing, Peking University; School of Computer Science, Peking University; Columbia University; YiXin-AILab, YIXIN; Southern University of Science and Technology(北京大学多媒体信息处理国家重点实验室; 北京大学计算机学院; 哥伦比亚大学; 壹芯人工智能实验室,壹芯; 南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SteerScope多维评估套件,用15项指标系统刻画23种干预方法在有效性与副作用间的权衡,发现激活干预尚未超越提示基线,且二者存在耦合。
AI 中文摘要
激活干预提供了一种轻量且灵活的方式来控制大语言模型(LLM)的行为。然而,有效的干预不仅需要诱导出预期行为:它还应限制非预期变化,并在不同输入和训练数据下保持稳健。现有评估仅零散地覆盖了这些维度。因此,有效性与副作用之间的权衡尚未被系统性地刻画。我们引入了SteerScope,一个双轴、多维度的评估套件,通过15项指标联合刻画干预结果和方法属性。我们评估目标有效性和副作用,涵盖语言质量、任务能力、安全性与可靠性,并进一步通过针对干预的样本效率和样本敏感性指标来评估泛化性和数据依赖性。我们并非在单一操作点上比较方法,而是刻画有效性与副作用之间的权衡。在匹配的模型、任务和评估协议下,我们对涵盖4个家族的23种方法进行了基准测试,包括提示、LoRA和SFT作为基线方法,并将该套件作为可扩展的代码库发布。我们发现,当前的激活干预方法在干预有效性与副作用的整体平衡上尚未超越提示干预基线:在两种模型规模下,没有一种被评估的激活干预方法能在不产生更大复合副作用的情况下实现更高的有效性。我们进一步发现了干预有效性与副作用之间的一致耦合。在分布外(OOD)提示下,目标有效性通常得以保持,而副作用往往变得更加显著,尤其是通过指令相关性和流畅性的下降表现出来。方法还表现出截然不同的样本效率特征。
英文摘要
Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result, the trade-offs between efficacy and side effects have not been systematically characterized. We introduce SteerScope, a two-axis, multi-dimensional evaluation suite that jointly characterizes steering outcomes and method properties through 15 metrics. We score target efficacy and side effects on language quality, task capabilities, and safety and reliability, and further assess generalization and data dependence through steering-specific metrics for sample efficiency and sample sensitivity. Rather than comparing methods at a single operating point, we characterize the trade-offs between efficacy and side effects. Under matched models, tasks, and evaluation protocols, we benchmark 23 methods spanning 4 families, including prompting, LoRA, and SFT as baseline methods, and release the suite as an extensible codebase. We find that current activation steering methods do not yet surpass the Prompt Steering baseline in their overall balance between steering efficacy and side effects: across both model scales, no evaluated activation steering method achieves higher efficacy without incurring greater composite side effects. We further uncover a consistent coupling between steering efficacy and side effects. Under OOD prompts, target efficacy is often preserved, whereas side effects tend to become more pronounced, particularly through declines in instruction relevance and fluency. Methods also exhibit sharply different sample-efficiency profiles.