发表机构
National Yang Ming Chiao Tung University; National Tsing Hua University(国立阳明交通大学; 国立清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对世界模型预测未来但缺乏可执行目标的问题,提出实体级目标读出接口,将终端目标作为显式输出,结合姿态预测与深度平移,在五个任务上平均成功率79.69%,零样本部署成功率达73.33%以上。
AI 中文摘要
生成式世界模型为操作场景如何向任务目标演变提供了丰富的预测,但这些未来状态并未直接暴露控制所需的紧凑任务变量。当仅以未来预测作为训练监督时,终端目标精度并非明确的学习目标,即使几何恢复可用。我们提出实体级目标读出(Entity-Level Goal Readout),一种学习得到的预测到执行接口,使可执行的终端目标成为3D轨迹世界模型的显式输出。该方法将对象中心姿态预测与基于观测深度的平移相结合,生成SE(3)中的紧凑目标。共享的姿态原生执行器(Pose-Native Executor)利用在线对象姿态反馈消费该固定目标,实现闭环控制,无需重新运行世界模型。在五个操作任务中,该流程平均成功率达79.69%。目标诊断直接测量终端目标精度,而受控平移扰动则刻画执行在目标误差下的退化情况。在Franka机械臂上的零样本部署在标称StackCube任务中达到73.33%的成功率,在存在干扰物时达66.67%,在策略训练中未见过的目标PickPlate任务中达75.00%。这些结果支持将预测到执行接口视为世界模型规划中显式的学习组件,而非控制流程中偶然的后处理。项目页面:此https URL
英文摘要
Generative world models provide rich predictions of how manipulation scenes may evolve toward task objectives, yet those futures do not directly expose the compact task variables required by control. When training supervises future prediction alone, terminal goal accuracy is not an explicit learning objective, even when geometric recovery is available. We present Entity-Level Goal Readout, a learned prediction-to-execution interface that makes the executable terminal goal an explicit output of a 3D trace world model. It combines object-centric pose prediction with translation grounded in observed depth to produce a compact goal in SE(3). A shared Pose-Native Executor consumes this fixed goal with online object-pose feedback for closed-loop control without rerunning the world model. Across five manipulation tasks, the pipeline achieves a mean success rate of 79.69%. Goal diagnostics directly measure terminal goal accuracy, while controlled translation perturbations characterize how execution degrades under goal error. Zero-shot deployment on a Franka arm achieves 73.33% success on nominal StackCube, 66.67% with distractors, and 75.00% on PickPlate with a target unseen during policy training. These results support treating the prediction-to-execution interface as an explicit learned component of world-model planning rather than incidental post-processing in the control pipeline itself. Project page: https://claire0730.github.io/executable-goals/
Comments9 pages, 8 figures, 4 tables. Project page: https://claire0730.github.io/executable-goals/ Code and models: https://github.com/Claire0730/executable-goals