MT-WAM:将单次前向预测表示重新定向用于动作生成
MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation
浏览论文内容
中文总结 AI 辅助
MT-WAM通过添加未来点轨迹和视觉特征监督,重新定向单次前向预测表示,提升动作生成性能,在LIBERO-Plus和真实任务上显著超越Fast-WAM。
中文摘要 AI 辅助
Fast-WAM表明,视频-动作联合训练无需在推理时生成未来视频即可提升控制性能,这使得来自单次视频扩散Transformer前向传播的表示成为动作生成的核心。然而,未来观测预测并未明确优先考虑控制所需的未来动态和视觉结构。我们提出MT-WAM,它保留原始训练目标,并添加针对未来二维点轨迹和视觉特征的互补监督。从视频主干网络最终块复制的轻量级双流分支提供目标特定处理,而结构化注意力掩码防止跨流注意力。运动流令牌为动作专家提供额外的动态条件。未来视觉特征预测在捕获物体和空间结构的特征空间中提供监督。这种监督训练视频主干网络在变化的视觉条件下为动作生成提供更具信息量的视觉上下文,而无需将视觉特征流令牌添加到动作条件中。在推理时,MT-WAM使用每次重新规划时计算一次的视频和运动缓存,并跳过未来视频预测。无需额外的具身策略预训练,MT-WAM在LIBERO上达到98.2%的成功率,在LIBERO-Plus上达到73.7%,后者超过Fast-WAM 23.8个百分点。在RoboTwin 2.0 Clean2Rand上,Random成功率从6.30%提升至19.40%;在四个真实世界任务中,平均成功率从67.0%提升至77.8%。
英文摘要
Fast-WAM shows that video-action co-training improves control without generating future video at inference, making the representation from a single video diffusion Transformer forward central to action generation. However, future-observation prediction does not explicitly prioritize the future dynamics and visual structure needed for control. We present MT-WAM, which retains the original training objectives and adds complementary supervision for future two-dimensional point trajectories and visual features. A lightweight dual-stream branch copied from the video backbone's final blocks provides target-specific processing, while a structured attention mask prevents cross-stream attention. Motion-stream tokens supply additional dynamics conditions to the action expert. Future visual-feature prediction provides supervision in a feature space that captures object and spatial structure. This supervision trains the video backbone to provide more informative visual context for action generation under changing visual conditions, without adding visual-feature-stream tokens to action conditioning. At inference, MT-WAM uses video and motion caches computed once per replan and skips future-video prediction. Without additional embodied policy pretraining, MT-WAM achieves 98.2% success on LIBERO and 73.7% on LIBERO-Plus, exceeding Fast-WAM by 23.8 percentage points on the latter. On RoboTwin 2.0 Clean2Rand, Random success increases from 6.30% to 19.40%; across four real-world tasks, average success increases from 67.0% to 77.8%.
发表机构
- Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。