发表机构
Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生成式机器人策略部署时的实时约束问题,提出INSPO方法,通过在输入噪声空间进行轨迹优化实现推理时引导,提升任务成功率与约束满足度,且运行时更低。
AI 中文摘要
生成式机器人策略能够表示多样化的多模态行为,但将预训练策略适配到部署时的约束(如避碰和朝向保持)仍具挑战性。现有的推理时引导方法通常通过迭代扩散或流过程应用梯度引导,这对于实时控制而言计算成本可能过高。我们提出INSPO,将单步生成策略的推理时引导形式化为策略输入噪声空间中的轨迹优化。通过在评估诱导状态轨迹上的约束的同时优化输入噪声,INSPO在不直接修改生成动作的情况下搜索策略诱导的行为空间。该优化包含一个正则化项,鼓励解与策略的输入分布保持一致,并通过基于群体的粒子优化在线求解。我们在基于状态和基于图像的任务特定策略以及通用视觉-语言-动作策略上评估了INSPO,涉及Push-T、Can拾放和LIBERO-Spatial任务。与最佳N采样和动作投影相比,INSPO提高了任务成功率和约束满足度,同时在更低运行时下与梯度引导生成相比具有竞争力。
英文摘要
Generative robot policies can represent diverse, multimodal behaviors, but adapting pretrained policies to deployment-time constraints such as collision avoidance and orientation maintenance remains challenging. Existing inference-time steering methods typically apply gradient guidance through iterative diffusion or flow processes, which can be computationally expensive for real-time control. We propose INSPO, which formulates inference-time steering of one-step generative policies as trajectory optimization in the policy's input noise space. By optimizing the input noise while evaluating constraints on the induced state trajectory, INSPO searches the policy-induced behavior space without directly modifying generated actions. The optimization includes a regularization term that encourages solutions to remain consistent with the policy's input distribution and is solved online using population-based particle optimization. We evaluate INSPO on state- and image-based task-specific policies and generalist vision-language-action policies across Push-T, Can pick-and-place, and LIBERO-Spatial. INSPO improves task success and constraint satisfaction over best-of-N sampling and action projection, while comparing favorably with gradient-guided generation at lower runtime.