发表机构
University of Arizona(亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出AIMES框架,利用因果激活引导和在线观察者反馈,在推理时自适应调整多价值干预强度,实现轻量级、状态感知的多值控制,优于固定联合引导。
AI 中文摘要
大型语言模型(LLMs)越来越多地被部署在需要响应多种可能相互影响的社会规范和人类价值观的场景中。激活引导提供了一种轻量级的替代基于训练的校准方法,通过在推理时修改内部激活来实现。然而,先前的人类价值引导方法大多孤立地考虑价值观,而直接组合多个方向依赖于固定的干预强度,无法响应模型不断演变的内部状态。基于这一关键观察,我们引入了AIMES,一个用于自适应多值激活引导的框架。AIMES为道德基础价值观构建了层特定的双极方向,并使用中间层词汇读出作为在线观察者。然后,一个由观察者引导的控制器在每个解码步骤根据当前观察到的状态调整每个请求的价值干预的强度,而无需训练单独的价值状态估计器。在多个指令调优模型家族、价值组合和干预深度中,我们发现多值可控性因价值组合和干预位置而异。与固定联合引导和基于提示的引导相比,AIMES显示出深度依赖的优势,这些优势在两个独立评估者中得到广泛支持,但在特定控制效果出现的精确深度上存在一些差异。这些优势伴随着比固定联合引导更小的实际激活空间干预和可比的响应质量。总体而言,我们的结果表明,在线观察者反馈可以为单次多值引导提供轻量级、状态感知的自适应。
英文摘要
Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model's evolving internal state. Motivated by this key observation, we introduce AIMES, a framework for adaptive multi-value activation steering. AIMES constructs layer-specific bipolar directions for moral-foundation values and uses intermediate-layer vocabulary readouts as online observers. An observer-guided controller then adapts the strength of each requested value intervention at every decoding step based on its current observed state, without training a separate value-state estimator. Across multiple instruction-tuned model families, value combinations, and intervention depths, we find that multi-value controllability varies across both value combinations and intervention locations. Compared with fixed joint steering and prompt-based steering, AIMES shows depth-dependent advantages that are broadly supported across two independent evaluators, with some variation in the precise depth at which specific control effects emerge. These advantages come with smaller realized activation-space interventions than fixed-joint steering and comparable response quality. Overall, our results suggest that online observer feedback can provide lightweight, state-aware adaptation for single-pass multi-value steering.