arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27033cs.RO

Riemann-1.0:面向物理人工智能的具身世界动作模型

Riemann-1.0: An Embodied World Action Model for Physical AI

  • Riemann Dynamics(黎曼动力公司)

机构由 AI 辅助整理,请以论文原文为准。

Haofeng Sun, Jiangbo Pei, Fei Kang, Zexiang Liu, Yaokun Li, Boyi Jiang, Hua Xue, Cindy Zhou, Wei Li, Yichen Wei, Mengyin An, Fanliang Zhao, Biao Jiang, Zile Wan… 展开作者

Haofeng Sun, Jiangbo Pei, Fei Kang, Zexiang Liu, Yaokun Li, Boyi Jiang, Hua Xue, Cindy Zhou, Wei Li, Yichen Wei, Mengyin An, Fanliang Zhao, Biao Jiang, Zile Wang, Yang Liu, Yangguang Li

AI总结:

本研究提出Riemann-1.0具身世界动作模型,通过统一建模与渐进式具身预训练,在多个模拟及真实机器人操作基准任务中取得最优性能,实现大规模具身经验向可泛化机器人操作能力的转化。

AI中文摘要:

我们提出Riemann-1.0,这是一款面向具身智能的全因果自回归世界动作模型。Riemann-1.0在统一的因果自回归序列中联合建模多视角视觉观测、机器人状态和具身特定动作,将机器人动作与世界演化表示为因果状态转移。与现有基于联合生成、视频优先预测或解耦建模范式的世界动作模型(WAM)不同,Riemann-1.0将在线机器人策略执行与动作条件化的世界模拟统一在单个模型中,使其既能作为可执行机器人策略,又能作为多具身视觉世界模拟器发挥作用。为了跨异构数据源扩展具身经验,我们进一步开发了渐进式具身预训练框架,该框架在统一的世界动作建模目标下,统一从自我中心人类视频、手持夹具演示和异构机器人轨迹中学习。基于20万小时以上的交互数据构建的Riemann-1.0,将大规模具身经验逐步转化为可执行的机器人操作能力。Riemann-1.0在模拟基准和真实世界操作任务中均取得了最先进的性能:在RoboTwin2.0上的成功率为94.3%,在LIBERO上为99.0%,在长程组合基准RoboCasa-365上为62.6%,比之前的最佳方法高出8.4%;在长程真实世界操作任务中,Riemann-1.0的成功率(SR)为85.0%,进展成功率(PSR)为94.4%,在SR上超过最强开源基准15%。这些结果表明,统一的世界动作建模结合渐进式具身预训练,可有效将大规模具身经验转化为可泛化的机器人操作能力。

英文摘要:

We introduce Riemann-1.0, a fully causal autoregressive World Action Model for embodied intelligence. Riemann-1.0 jointly models multi-view visual observations, robot states, and embodiment-specific actions within a unified causal autoregressive sequence, representing robot actions and world evolution as causal state transitions. Unlike existing WAMs based on joint generation, video-first prediction, or decoupled modeling paradigms, Riemann-1.0 unifies online robot policy execution and action-conditioned world simulation within a single model, enabling it to function as both an executable robot policy and a multi-embodiment visual world simulator. To scale embodied experience across heterogeneous data sources, we further develop a progressive embodied pretraining framework that unifies learning from egocentric human videos, handheld-gripper demonstrations, and heterogeneous robot trajectories under a shared World Action Modeling objective. Built upon 200K+ hours of interaction data, Riemann-1.0 progressively transfers large-scale embodied experience into executable robot manipulation capabilities. Riemann-1.0 achieves state-of-the-art performance across both simulation benchmarks and real-world manipulation tasks. It achieves success rates of 94.3% on RoboTwin2.0, 99.0% on LIBERO, and 62.6% on the long-horizon compositional benchmark RoboCasa-365, outperforming the previous best method by 8.4% On long-horizon real-world manipulation tasks, Riemann-1.0 achieves a Success Rate (SR) of 85.0% and a Progress Success Rate (PSR) of 94.4%, exceeding the strongest open-source baseline by 15% in SR. These results demonstrate that unified World Action Modeling together with progressive embodied pretraining effectively transforms large-scale embodied experience into generalizable robot manipulation capabilities.

↑