arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ForeAct3D:面向VLA策略的策略接地未来世界建模

ForeAct3D: Policy-Grounded Future World Modeling for VLA Policies

Zhe Tao, Feiran Wang, Gaowen Liu, Ramana Rao Kompella$, Yan Yan

arXiv 2610.04607首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; University of Illinois Chicago; Cisco(伊利诺伊大学厄巴纳-香槟分校; 伊利诺伊大学芝加哥分校; 思科)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ForeAct3D通过策略接地未来世界建模,利用可学习几何查询和物理一致性约束,在VLA策略中实现语义3D场景预测,无需推理时预测,显著提升LIBERO、CALVIN及真实操作任务性能。

AI 中文摘要

机器人需要预测其动作将如何改变世界,因为操作的成功取决于由此产生的接触和物体运动。然而,现有的视觉-语言-动作(VLA)策略从共享特征预测未来观测,使得预测与策略实际执行的行动脱节,并且对场景可能如何演变没有施加物理约束。我们提出了ForeAct3D,一个在VLA策略中进行策略接地未来世界建模的框架。可学习的几何查询将策略表示中的深度、语义分割和相机姿态解码为当前和未来的语义3D场景状态,并且未来查询以策略生成的动作块为条件,将预测基于计划的交互。物理一致性闭合通过背景静态性和实例级刚性关联两种状态,并将腕部相机姿态锚定到末端执行器运动学。这些目标在训练期间塑造用于动作生成的共享表示,并且在推理时不需要未来预测。在没有机器人预训练的情况下,ForeAct3D在LIBERO上实现了98.3%的平均成功率,在CALVIN上实现了3.73的平均任务长度,在每个套件上都优于其基础策略。消融实验表明,语义3D监督、物理一致性和动作条件各自提高了操作性能,并且动作条件显著改善了未来物体定位。在空间放置、物体插入和顺序操作上的真实世界实验进一步将平均成功率从6.7%提高到37.8%,超过了基础策略。项目页面和代码可在该https URL获取。

英文摘要

Robots need to anticipate how their actions will change the world, since manipulation success hinges on the resulting contacts and object motions. However, existing Vision-Language-Action (VLA) policies that predict future observations from shared features leave the forecast decoupled from the actions the policy will actually execute, and impose no physical constraints on how the scene may evolve. We introduce ForeAct3D, a framework for policy-grounded future world modeling within VLA policies. Learnable geometric queries decode depth, semantic segmentation, and camera pose from the policy representation into current and future semantic 3D scene states, and the future queries are conditioned on the policy-generated action chunk to ground the forecast in the planned interaction. A physical-consistency closure relates the two states through background staticity and instance-level rigidity, and anchors the wrist-camera pose to end-effector kinematics. These objectives shape the shared representation used for action generation during training, and no future prediction is required at inference. Without robot pretraining, ForeAct3D achieves 98.3\% average success on LIBERO and an average task length of 3.73 on CALVIN, outperforming its base policy on every suite. Ablations show that semantic 3D supervision, physical consistency, and action conditioning each improve manipulation performance, and that action conditioning substantially improves future object localization. Real-world experiments on spatial placement, object insertion, and sequential manipulation further raise average success from 6.7\% to 37.8\% over the base policy. The project page and code are available at https://github.com/anthonytao80-crypto/ForeAct3D.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑