arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KineWorld:用于具身世界建模的动作诱导传输场

KineWorld: Action-Induced Transport Fields for Embodied World Modeling

Ziying Song, Yuchen Liu, Zhuoran Xu, Ziyang Liu, Jian Jin, Jiangtao Su, Haibao Yu, Lei Yang, Yuanpei Chen

arXiv 2610.06349首次发表:更新:

发表机构

Nanyang Technological University; North University of China; Psibot; The Hong Kong Polytechnic University; China Academy of Information and Communications Technology; University of Hong Kong; Peking University(南洋理工大学; 中北大学; Psibot; 香港理工大学; 中国信息通信研究院; 香港大学; 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KineWorld通过运动传输提升和传输感知世界扩散,利用机器人运动学构建传输场并重加权生成目标,在RoboTwin 2.0数据上显著提升具身世界模型的预测性能。

AI 中文摘要

具身世界模型在执行前预测候选动作的视觉后果。然而,现有的动作条件世界模型通常采用均匀加权的视觉生成目标,这可能与具身预测需求不一致。即使有显式的运动条件,这些目标也可能低估对交互至关重要的空间稀疏变化。我们提出KineWorld,一个传输感知的世界建模框架,将机器人运动学从运动条件扩展到生成性监督的空间分配。运动传输提升(KTL)从命令的机器人运动构建渲染器派生、相机对齐的传输场。传输感知世界扩散(TAWD)在视频潜在网格上校准其运动支持,并通过均匀和传输聚焦分布的归一化混合重新加权未来RGB流匹配。我们使用来自RoboTwin 2.0的ALOHA-AgileX双臂操作数据训练KineWorld。KineWorld在单视图评估中达到EWMScore-P 68.95,在多视图评估中达到TWB-Score 54.82。这些结果支持从外观拟合向动作后果建模的转变,用于具身决策。

英文摘要

Embodied world models predict the visual consequences of candidate actions before execution. However, existing action-conditioned world models often adopt uniformly weighted visual generation objectives that can be misaligned with embodied prediction needs. Even with explicit motion conditioning, these objectives can underemphasize spatially sparse changes that are critical to interaction. We propose KineWorld, a transport-aware world-modeling framework that extends robot kinematics from motion conditioning to the spatial allocation of generative supervision. Kinematic Transport Lifting (KTL) constructs renderer-derived, camera-aligned transport fields from commanded robot motion. Transport-Aware World Diffusion (TAWD) calibrates their motion support on the video-latent grid and reweights future-RGB flow matching through a normalized mixture of uniform and transport-focused distributions. We train KineWorld using ALOHA-AgileX bimanual manipulation data from RoboTwin 2.0. KineWorld achieves an EWMScore-P of 68.95 in single-view evaluation and a TWB-Score of 54.82 in multi-view evaluation. These results support a shift from appearance fitting toward action-consequence modeling for embodied decision-making.

Comments36 pages. Project page and code: https://modaxiansheng.github.io/KineWorld/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑