arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GigaWorld-Policy-0.5:由自动研究赋能的更快更强的世界行动模型

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

GigaWorld Team, Angen Ye, Angyuan Ma, Boyuan Wang, Chaojun Ni, Fangzheng Ye, Guan Huang, Guo Li, Guosheng Zhao, Haodong Yan, Hengtao Li, Jiwen Lu, Kai Wang, Mingming Yu, Qitang Hu, Qiuping Deng, Songling Liu, Xiaoyu Tian, Xiaofeng Wang, Xinyu Zhou, Xiuwei Xu, Xinze Chen, Yang Wang, Yejun Zeng, Yifan Chang, Yun Ye, Zhenyu Wu, Zhanqian Wu, Zheng Zhu

arXiv 2607.13960首次发表:更新:

发表机构

GigaAI; Tsinghua University(极佳科技; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在改进机器人策略学习,提出GigaWorld-Policy-0.5这一增强型以动作中心的WAM。预训练采用混合策略加强视觉与动作耦合,推理引入新架构提升效率,还用自动研究管道搜索训练配置,有效提升了机器人控制推理效率并保留训练优势。

AI 中文摘要

世界行动模型(WAMs)通过联合对动作和未来视觉观察进行建模,利用未来场景演变作为物理基础动作生成的密集监督,来改进机器人策略学习。然而,现有WAMs的常见设计是在推理时显式生成未来视频,这会带来大量计算开销并阻碍实时闭环部署。GigaWorld-Policy以以动作中心的公式解决此问题,训练时使用未来视觉动态,推理时仅使用动作解码。在此框架基础上,提出了GigaWorld-Policy-0.5,一种为更高效机器人控制设计的增强型以动作中心的WAM。预训练时,采用混合动作条件世界建模(AC-WM)和WAM训练策略,加强视觉动态与机器人动作之间的耦合,提高动作表示对下游策略学习的可迁移性。为实现高效推理,引入了Transformer混合架构,将视觉动态建模和动作生成分离为专门的专家,减少仅动作推理期间的主动计算,在本地RTX 4090设置上实现85毫秒的推理延迟。此外,采用基于代理的自动研究管道系统地搜索训练配置,更高效地识别最佳实验设置,减少超参数调整所需的时间和人工干预。实验和消融表明,GigaWorld-Policy-0.5在提高机器人控制推理效率的同时,保留了未来视觉动态的训练优势。

英文摘要

World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly generate future videos at inference time, incurring substantial computational overhead and hindering real-time closed-loop deployment. GigaWorld-Policy addresses this issue with an action-centered formulation, where future visual dynamics are used during training while action-only decoding is used at inference time. Building upon this framework, we present GigaWorld-Policy-0.5, an enhanced action-centered WAM designed for more efficient robot control. During pretraining, GigaWorld-Policy-0.5 adopts a mixed Action-Conditioned World Modeling (AC-WM) and WAM training strategy. This strengthens the coupling between visual dynamics and robot actions and improves the transferability of action representations for downstream policy learning. For efficient inference, GigaWorld-Policy-0.5 introduces a Mixture-of-Transformers architecture that separates visual dynamics modeling and action generation into specialized experts, reducing active computation during action-only inference and achieving 85 ms inference latency on a local RTX 4090 setup. In addition, we employ an agent-based AutoResearch pipeline to systematically search training configurations, enabling more efficient identification of optimal experimental setups while reducing the time and manual intervention required for hyperparameter tuning. Experiments and ablations show that GigaWorld-Policy-0.5 preserves the training benefits of future visual dynamics while improving inference efficiency for robot control.

Commentsproject page: https://open-gigaai.github.io/giga-world-policy/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑