arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33172cs.ROcs.AIcs.CVcs.LG

通过反事实规划的世界-动作模型动态操作

Dynamic Manipulation with World-Action Models via Counterfactual Planning

Sunwoo Park, Wonbin Lee, Seonghyun Jin, Youngmin Kim, Jangho Park, Jong Chul Ye

AI总结:

针对世界-动作模型在静态演示训练后难以操作移动目标的问题,提出动态预测规划框架,通过反事实规划解耦计划生成与执行状态,利用预测展开估计交互时间并结合目标运动预测未来位置,构建反事实观察以调用已有技能,实现无需额外动态数据训练的实时动态操作,在仿真和真实机器人上均优于基线。

AI中文摘要:

在静态演示上训练的世界-动作模型(WAMs)即使具备所需的操作技能,也常常无法操作移动目标。我们将这种失败归因于目标-响应崩溃:随着执行的推进,策略越来越偏向于其正在进行的行为的学习延续,而对目标重新定位的响应性降低。为了弥合模型已学内容与从当前上下文可生成内容之间的差距,我们将动态操作表述为反事实规划,通过将用于计划生成的上下文与用于执行的物理状态解耦。我们的框架,动态预测规划(DPP),首先利用WAM的预测展开来估计交互预期发生的时间,并将该时间估计与观察到的目标运动相结合,以预测目标的未来交互位置。DPP随后构建一个反事实观察,将该预测的目标位置置于熟悉的机器人上下文中,使模型能够调用现有的操作技能,而不是从不熟悉的机器人-目标配置中生成恢复行为。生成的计划在执行过程中连接到机器人的实际状态。DPP使得在单个消费级GPU上无需额外动态数据训练即可实现实时动态操作。在仿真和真实机器人上的实验表明,在多样化的目标运动下均有一致的改进,仿真性能超过了所有评估的基线,包括额外在动态数据上训练的方法。项目页面:此 https URL

英文摘要:

World-Action models (WAMs) trained on static demonstrations often fail to manipulate moving targets even when they possess the required manipulation skills. We attribute this failure to target-response collapse: as execution advances, the policy becomes increasingly biased toward the learned continuation of its ongoing behavior and less responsive to target relocation. To bridge the gap between what the model has learned and what it can generate from the current context, we formulate dynamic manipulation as counterfactual planning by decoupling the context used for plan generation from the physical state used for execution. Our framework, Dynamic Predictive Planning (DPP), first uses the WAM's predictive rollout to estimate when an interaction is expected to occur, and combines this timing estimate with observed target motion to predict the target's future interaction position. DPP then constructs a counterfactual observation that places this predicted target position in a familiar robot context, allowing the model to invoke an existing manipulation skill rather than generate a recovery behavior from an unfamiliar robot-target configuration. The resulting plan is connected to the robot's actual state during execution. DPP enables real-time dynamic manipulation on a single consumer GPU without additional training on dynamic data. Experiments in simulation and on a real robot demonstrate consistent improvements across diverse target motions, with simulation performance surpassing all evaluated baselines, including methods additionally trained on dynamic data. Project page: https://methoder00.github.io/DPP/

补充信息

↑