AI 中文总结
研究如何让基于RGB的世界-动作模型从外观主导的重建转向交互诱导的视觉动态,提出DC-WAM框架,通过重新分配监督和计算,结合时间差分流匹配等方法,提升策略性能,尤其在多种分布外扰动下表现出色。
AI 中文摘要
世界-动作模型(WAMs)通过未来视觉预测增强机器人策略,但视觉模态应学习什么用于控制尚不清楚。逼真的未来预测虽提供密集监督,但计算量大且会将能力分配到与动作选择弱相关的纹理、光照和背景变化上。近期高效的WAM变体表明视频分支的主要好处可能不在渲染的未来本身,而在训练期间诱导的与控制相关的视觉表示。本文从以动态为中心的视角重新审视未来视频预测,提出DC-WAM,一个在RGB视频分支中重新分配监督和计算的以动态为中心的WAM框架。在监督层面,DC-WAM将时间差分流匹配与轨迹引导加权相结合,强调密集的时间变化和抓取器、被操作物体及接触区域移动的局部区域。在推理层面,DynaRoute预测令牌级的动态相关性并将其转换为注意力偏差,引导模型关注与控制相关的未来令牌。仿真和实际操作任务实验表明,DC-WAM持续提升策略性能,尤其在光照、物体外观和背景纹理的分布外扰动下。
英文摘要
World-Action Models (WAMs) augment robot policies with future visual prediction, but it remains unclear what the visual modality should learn for control. While photorealistic future prediction provides dense supervision, it also incurs substantial computation and can allocate capacity to texture, illumination, and background variations that are only weakly related to action selection. Recent efficient WAM variants suggest that the main benefit of the video branch may not lie in the rendered future itself, but in the control-relevant visual representations induced during training. In this work, we revisit future video prediction from a dynamic-centric perspective and ask whether an existing RGB-based WAM can be redirected from appearance-dominated reconstruction toward interaction-induced visual dynamics without introducing additional modality-specific predictions or online inputs at deployment. We propose DC-WAM, a dynamic-centric WAM framework that redistributes supervision and computation in the RGB video branch. At the supervision level, DC-WAM combines temporal-difference flow matching with trajectory-guided weighting, emphasizing dense temporal changes and localized regions where the gripper, manipulated objects, and contact areas move. At the reasoning level, DynaRoute predicts token-wise dynamic relevance and converts it into an attention bias, guiding the model toward control-relevant future tokens. Experiments in simulation and on real-world manipulation tasks show that DC-WAM consistently improves policy performance, especially under out-of-distribution perturbations in lighting, object appearance, and background texture.