FoMoVLA:为视觉-语言-动作模型架起视觉预见与运动引导的桥梁
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
AI总结:
研究针对视觉-语言-动作模型缺乏对世界动态向前预测及运动引导不足问题,提出FoMoVLA框架,联合学习未来特征预见与稀疏点跟踪增强VLA表示,经实验验证其性能先进且泛化能力强。
AI中文摘要:
视觉-语言-动作(VLA)模型在视觉运动策略学习中取得了显著成果,但本质上仍具有反应性,将当前观察和语言映射到动作时未对世界动态进行明确的向前预测。现有视觉预见方法能预测未来视觉状态,但缺乏明确的运动引导。本文认为未来特征预测和稀疏点跟踪是自然互补的,前者提供目标状态,后者捕捉通向该目标的连续运动路径。为此提出FoMoVLA框架,通过联合学习未来特征预见和稀疏二维点跟踪,用明确的时空监督增强VLA表示,提升连续动作策略。该框架引入紧凑预见令牌解码未来特征状态,解码稀疏时间二维点轨迹以建模紧凑几何运动,并通过轻量级未来条件交叉注意力模块将两者结合,实现预期状态和点动态之间的一致推理。在LIBERO、RoboCasa GR-1 Tabletop和LIBERO-Plus上的大量实验证明了其具有最先进的性能和强大的零样本泛化能力。
英文摘要:
Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods predict future visual states but lack explicit motion guidance: they show where to go but not how to get there. We argue that future feature prediction and sparse point tracking are naturally complementary: the former provides the goal state, while the latter captures the continuous motion path toward it. We propose FoMoVLA, a framework that augments VLA representations with explicit spatio-temporal supervision by jointly learning future feature foresight and sparse 2D point tracking, enhancing the continuous action policy. FoMoVLA introduces compact foresight tokens to decode future feature states, decodes sparse temporal 2D point trajectories to model compact geometric motion, and couples both through a lightweight future-conditioned cross-attention module that enables consistent reasoning between anticipated states and point dynamics. Extensive experiments on LIBERO, RoboCasa GR-1 Tabletop, and LIBERO-Plus demonstrate state-of-the-art performance and strong zero-shot generalization. Project page is available at https://liauto-research.github.io/FoMoVLA.