AI 中文总结
针对前沿大语言模型常违反指令层级的问题,提出无需训练的V-Steer方法,通过编辑缓存值向量调控指令层级,提升角色冲突基准准确率,性能优于仅提示基线,部分场景达到SoTA训练方法水平。
AI 中文摘要
指令层级是语言模型部署的核心安全假设:更高优先级的输入(如系统提示)应覆盖用户或工具的低优先级冲突输入,但前沿大语言模型(LLM)常违反该层级。我们提出V-Steer,一种无需训练的推理时方法,通过编辑提示位置的缓存值向量来恢复特权影响。利用首个下一个词预测的直接logit归因,V-Steer识别低优先级片段主导特权片段的注意力头,随后通过对缓存V张量的原位乘法编辑,增强特权片段并抑制冲突的低优先级片段。由于该方法仅作用于缓存值,它与融合注意力后端兼容,且仅增加一次预填充开销。在7B至70B的模型上,该归因引导的干预将受控角色冲突基准的主要约束准确率从低于18%提升至92%;在更广泛的指令层级评估中,其显著优于仅提示的基线方法,且在4种LLM规模中的3种上,达到或超越了基于训练的当前最优(SoTA)方法,解码速度开销可忽略不计。代码可在指定URL获取。
英文摘要
Instruction hierarchies are a core safety assumption of language model deployment: higher priority inputs, such as system prompts, should override conflicting lower priority inputs from users or tools. Yet frontier LLMs often violate this hierarchy. We introduce V-Steer, a training-free inference time method that restores privileged influence by editing cached value vectors at prompt positions. Using direct logit attribution on the first next token prediction, V-Steer identifies heads where lower priority spans dominate privileged ones, then boosts privileged spans and suppresses conflicting lower priority spans through in-place multiplicative edits to cached V tensors. Since the method acts only on cached values, it remains compatible with fused attention backends and adds only a one time prefill overhead. Across models from 7B to 70B, this attribution guided intervention raises primary constraint accuracy from under 18% up to 92% on controlled role conflict benchmarks, and on broader instruction hierarchy evaluations substantially outperforms prompt only baselines while matching or exceeding SoTA training based methods on 3 of 4 scales of LLMs, with negligible decoding-speed overhead. The code is available at https://github.com/cindy2000sh/v-steer.
CommentsPublished as a conference paper at COLM '26; the first two authors contributed equally to the work. 24 pages, 9 figures