AI 中文总结
针对现有激活引导无法引导双路径网格世界简单策略的问题,提出PGS方法,经网格世界、国际象棋、竞技足球实验验证其校准性、可逆性及行为适配的可组合性与跨域迁移性。
AI 中文摘要
激活引导已在大型语言模型中成为一种轻量级替代方案,用于在推理时动态改变模型的行为。然而,我们证明现有引导方法甚至无法在双路径网格世界环境中引导一个简单策略。为解决这一局限,我们提出策略梯度引导(Policy Gradient Steering, PGS),将引导建模为强化学习问题。PGS通过在少量回合或演示中累积临时行为目标的梯度,构建可移除的任务向量。我们首先在双路径网格世界环境中验证PGS的校准性与可逆性;随后使用国际象棋谜题,分别评估独立拟合的PGS向量单独及组合使用的效果,发现兼容的战术目标会建设性地累积;最后在竞技足球场景中,证明PGS可改变特定团队行为,且其效果可跨对手迁移。综上,这些结果表明策略梯度为在不同决策领域构建临时且可组合的行为适配提供了自然接口。
英文摘要
Activation steering has emerged in large language models as a lightweight alternative for dynamically changing a model's behavior at inference time. However, we show that existing steering methods fail to steer even a simple policy in a two-route gridworld environment. To address this limitation, we propose Policy Gradient Steering (PGS), which formulates steering as a reinforcement learning problem. PGS accumulates gradients of a temporary behavioral objective over a small set of rollouts or demonstrations to construct a removable task vector. We first demonstrate the calibration and reversibility of PGS in a two-route gridworld environment. Using chess puzzles, we then evaluate independently fitted PGS vectors both in isolation and in combination, finding that compatible tactical objectives accumulate constructively. Finally, in competitive football, we show that PGS can alter specific team behaviors and that its effects transfer across opponents. Together, these results show that policy gradients provide a natural interface for constructing temporary and composable behavioral adaptations across diverse decision-making domains.