arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39324cs.ROcs.CVcs.LG

MotionWeave:面向视觉-语言-动作策略的运动中心未来动态学习

MotionWeave: Learning Motion-Centered Future Dynamics for Vision-Language-Action Policies

Jingqiu Wang, Yan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出MotionWeave框架,通过动作诱导运动定位与水平残差合成,实现运动中心未来动态建模,在MetaWorld任务上平均成功率75.3%,较π0提升8.6%。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型近期引入世界模型,以提供超越稀疏动作标签的更丰富的动态监督。然而,显式预测未来图像或视频可能包含与控制无关的外观信息,而基于整体未来视觉表征和共享全局动作特征生成的引导可能无法建立动作与局部视觉变化之间的时间步对应关系。为解决此问题,我们提出MotionWeave,一种面向动作块预测的运动中心未来动态框架,包含两个模块:动作诱导运动定位器(AIMG)和水平残差合成器(HRC)。具体而言,AIMG基于动作和本体感受表征构建水平特定查询,从当前视觉标记中定位与每个未来动作时间步相关的交互区域。HRC提取相邻水平交互表征之间的差异,将其编码为时间运动线索,并通过门控残差注入动作标记。训练期间,利用从未来帧渲染的机械臂掩码构建基于KL散度的运动定位监督,而推理时仅使用当前观测。在六个MetaWorld任务上,MotionWeave实现了75.3%的平均成功率,相比π0(66.7%)绝对提升8.6%,尤其在持续交互任务上表现突出。我们的代码可在以下网址获取:此https URL。

英文摘要

Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. However, explicitly predicting future images or videos may include control-irrelevant appearance, while guidance derived from holistic future visual representations and shared global action features may fail to establish timestep-specific correspondence between actions and local visual changes. To address this issue, we propose MotionWeave, a motion-centric future-dynamics framework for action-chunk prediction with two modules: the Action-Induced Motion Grounder (AIMG) and the Horizon Residual Composer (HRC). Specifically, AIMG conditions on action and proprioceptive representations to construct horizon-specific queries that localize interaction regions associated with each future action timestep from current visual tokens. HRC extracts differences between interaction representations at adjacent horizons, encodes them as temporal motion cues, and injects them into action tokens through a gated residual. During training, robot-arm masks rendered from future frames are used to construct KL-based motion-grounding supervision, while inference uses only the current observation. On six MetaWorld tasks, MotionWeave achieves a 75.3% average success rate, an absolute gain of 8.6% over π0 (66.7%), especially on sustained-interaction tasks. Our code is available at https://github.com/autu-mn/MotionWeave.

发表机构

  • East China Normal University(华东师范大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑