arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlowWAM:光流作为世界动作模型的统一动作表示

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Yixiang Chen, Peiyan Li, Yuan Xu, Qisen Ma, Jiabing Yang, Kai Wang, Jianhua Yang, Dong An, He Guan, Gaoteng Liu, Jianlou Si, Jun Huang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang

arXiv 2607.13017首次发表:更新:

发表机构

New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; FiveAges; MBZUAI; Alibaba Group(中国科学院自动化研究所模式识别国家重点实验室; 中国科学院大学人工智能学院; 无; 无; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对世界动作模型控制中动作表示难题,提出FlowWAM双流扩散框架,以光流为统一动作表示。该框架可实现WAMs两种模式,能利用无动作标签视频预训练,实验表明在操纵和世界建模任务中表现优于基线。

AI 中文摘要

世界动作模型(WAMs)可利用预训练视频生成器进行世界建模和动作预测。但直接用于控制面临挑战:如何以合适形式表示动作,既与预训练视频生成器匹配,又携带足够运动线索用于精确控制。现有数值动作和视觉动作表示均有不足。我们提出FlowWAM,一种双流扩散框架,采用光流作为统一的、视频原生的动作表示。通过在共享预训练视频生成器中联合建模,FlowWAM可实现WAMs的两种模式。在策略模式下用于动作预测,在世界模型模式下用目标流序列指导未来视频生成。此外,光流可从无动作标签的原始视频中轻松提取,能利用大规模无动作标签视频数据集进行预训练。实验表明,基于光流的动作表示在两种模式下均有提升。在RoboTwin操纵任务中,在Clean设置下成功率达92.94%,在Random设置下为92.14%,优于VLA和WAM基线。在WorldArena世界建模任务中,实现最佳总体EWMScore(63.71),轨迹精度相对提高18.4%。更多结果可在项目网站查看。

英文摘要

World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control. Existing numerical actions fail to satisfy the former, and prior visual action representations overlook the temporal motion structure across frames. We address this issue with FlowWAM, a dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Flow videos share the same format as RGB videos and encode rich per-pixel displacement. By jointly modeling them within a shared pretrained video generator, FlowWAM can naturally implement two modes of WAMs. In policy mode, FlowWAM generates flow for action prediction, while in world-model mode, it uses target flow sequences to guide future video generation. Moreover, since flow can be easily extracted from raw videos without action labels, FlowWAM can leverage large-scale action-unlabeled video datasets for pretraining. We empirically find that our flow-based action representation delivers gains across both modes. On RoboTwin manipulation, FlowWAM raises the success rate to 92.94% on the Clean setting and 92.14% on Random, outperforming both VLA and WAM baselines. On WorldArena world modeling, it achieves the best overall EWMScore (63.71) with an 18.4% relative improvement in trajectory accuracy. More results can be found on our project website: https://flow-wam.github.io .

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑