发表机构
Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出在推理时通过动态屏障引导和级联过滤,使冻结的生成式机器人策略遵守词典序部署偏好,在导航和操作任务中提升符合性且不损害成功率。
AI 中文摘要
预训练的生成式机器人策略能够在多种环境中产生有效的行为,但在部署时可能会遇到训练中未表示的约束和偏好。此外,在部署时,操作员、用户或应用可能会为这些约束和偏好分配优先级顺序,且该顺序可能因部署而异。例如,特定于具体形态的可行性约束可能需要首先满足,而用户特定的偏好则指导在可行选项中的行为选择。我们证明,一个冻结的生成式机器人策略——基于扩散或流匹配——可以在推理时被引导以尊重此类词典序排列的部署目标。为实现这一点,我们对采样器引入了两项修改。首先,我们将动态屏障引导应用于采样的轨迹,限制低优先级的更新,使得高优先级的代价不会增加(达到一阶)。其次,我们使用一个级联来选择执行的样本,该级联根据每个优先级级别依次过滤候选样本。策略权重保持不变。在一个导航基准上,我们证明了我们的方法在成功率、可穿越性和偏好符合性方面优于冻结策略,并且比调优的加权和基线实现了显著更好的符合性。相同的方法迁移到LIBERO上的流匹配操作策略,在提高符合性的同时不降低任务成功率。一项受控的操作研究进一步表明,在固定权重能够匹配所需顺序的设置中,动态屏障在更广泛的参数设置范围内达到了可比的最佳性能。
英文摘要
Pretrained generative robot policies can produce effective behaviors across diverse environments, but deployment can lead to requirements and preferences that may not have been represented during training. Furthermore, at deployment, an operator, user, or application may assign these requirements and preferences a priority order that can vary across deployments. For example, embodiment-specific feasibility constraints may need to be satisfied first, while user-specific preferences guide behavior among the feasible options. We show that a frozen generative robot policy---based on either diffusion or flow matching---can be steered at inference time to respect such lexicographically ordered deployment objectives. To achieve this, we introduce two modifications to the sampler. First, we apply dynamic-barrier guidance to sampled trajectories, constraining lower-priority updates so that higher-priority costs do not increase (up to first order). Second, we select the executed sample using a cascade that successively filters candidate samples according to each priority level. The policy weights remain unchanged. On a navigation benchmark, we demonstrate that our method improves success, traversability, and preference compliance over the frozen policy, and achieves substantially better compliance than tuned weighted-sum baselines. The same method transfers to a flow-matching manipulation policy on LIBERO, where it improves compliance without reducing task success. A controlled manipulation study further shows that, in settings where a fixed weight can match the desired ordering, the dynamic barrier reaches comparable best performance over a substantially wider range of parameter settings.
Comments11 pages