发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
代理策略引导通过训练两个轻量级代理策略的校准速度差,在推理时引导冻结基础策略,实现任务专用化,提升成功率53%且保持广泛能力。
AI 中文摘要
通用机器人策略从大规模数据中携带了广泛的操作先验,但将它们专门化到新任务仍然是部署的瓶颈。这需要在有限的演示中引出任务特定的行为,同时不削弱其广泛的能力。我们引入了代理策略引导(Proxy Policy Steering, PPS),一种推理时自适应方法,通过训练两个轻量级代理策略来解决这一挑战,其校准的速度空间差异引导冻结的基础采样器。参考代理在目标任务观测上建模冻结基础的行为,而任务代理(从参考初始化)捕获该行为在任务监督下的变化。它们的差异形成一个校准的速度空间残差,在每个去噪步骤中引导冻结的基础采样器。我们确定了该残差隔离任务监督引起的变化的条件,并进行了实证验证。由于基础从未被直接修改,其广泛的能力在推理时仍然可用,包括演示本身未涉及的行为,如失败恢复。自适应仅需要基础的向前速度预测,使得PPS训练轻量级,并且即使无法访问基础参数也能应用。在8个真实世界和4个模拟操作任务上,PPS将最先进的pi 0.5基础策略的平均绝对成功率提升了53%,在基础从未解决的任务上实现了从零到一的提升,同时保持了基础的广泛能力。PPS优于LoRA微调、从零训练的专家、残差策略以及先前的推理时引导方法。
英文摘要
Generalist robot policies carry broad manipulation priors from large-scale data, but specializing them to a new task remains the deployment bottleneck. This requires eliciting task-specific behavior from limited demonstrations without degrading their broad capabilities. We introduce Proxy Policy Steering (PPS), an inference-time adaptation method that resolves this challenge by training two lightweight proxy policies whose calibrated velocity-space difference steers the frozen base sampler. A reference proxy models the frozen base's behavior on target-task observations, and a task proxy, initialized from the reference, captures how this behavior changes under task supervision. Their difference forms a calibrated velocity-space residual that steers the frozen base sampler at every denoising step. We identify the conditions under which this residual isolates the change induced by task supervision, and validate them empirically. Because the base is never directly modified, its broad capabilities remain available at inference, including behaviors such as recovery from failure that the demonstrations themselves do not exercise. Adaptation requires only forward velocity predictions from the base, making PPS lightweight to train and applicable even without access to the base's parameters. On 8 real-world and 4 simulation manipulation tasks, PPS lifts the state-of-the-art pi 0.5 base policy by 53% absolute success rate on average, with zero-to-one gains on tasks the base never solves, while preserving the base's broad capabilities. PPS outperforms LoRA fine-tuning, from-scratch specialists, residual policies, and prior inference-time steering methods.