Riemann-1.0:面向物理人工智能的具身世界动作模型
Riemann-1.0: An Embodied World Action Model for Physical AI
- Riemann Dynamics(黎曼动力公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出Riemann-1.0具身世界动作模型,通过统一建模与渐进式具身预训练,在多个模拟及真实机器人操作基准任务中取得最优性能,实现大规模具身经验向可泛化机器人操作能力的转化。
AI中文摘要:
我们提出Riemann-1.0,这是一款面向具身智能的全因果自回归世界动作模型。Riemann-1.0在统一的因果自回归序列中联合建模多视角视觉观测、机器人状态和具身特定动作,将机器人动作与世界演化表示为因果状态转移。与现有基于联合生成、视频优先预测或解耦建模范式的世界动作模型(WAM)不同,Riemann-1.0将在线机器人策略执行与动作条件化的世界模拟统一在单个模型中,使其既能作为可执行机器人策略,又能作为多具身视觉世界模拟器发挥作用。为了跨异构数据源扩展具身经验,我们进一步开发了渐进式具身预训练框架,该框架在统一的世界动作建模目标下,统一从自我中心人类视频、手持夹具演示和异构机器人轨迹中学习。基于20万小时以上的交互数据构建的Riemann-1.0,将大规模具身经验逐步转化为可执行的机器人操作能力。Riemann-1.0在模拟基准和真实世界操作任务中均取得了最先进的性能:在RoboTwin2.0上的成功率为94.3%,在LIBERO上为99.0%,在长程组合基准RoboCasa-365上为62.6%,比之前的最佳方法高出8.4%;在长程真实世界操作任务中,Riemann-1.0的成功率(SR)为85.0%,进展成功率(PSR)为94.4%,在SR上超过最强开源基准15%。这些结果表明,统一的世界动作建模结合渐进式具身预训练,可有效将大规模具身经验转化为可泛化的机器人操作能力。
英文摘要:
We introduce Riemann-1.0, a fully causal autoregressive World Action Model for embodied intelligence. Riemann-1.0 jointly models multi-view visual observations, robot states, and embodiment-specific actions within a unified causal autoregressive sequence, representing robot actions and world evolution as causal state transitions. Unlike existing WAMs based on joint generation, video-first prediction, or decoupled modeling paradigms, Riemann-1.0 unifies online robot policy execution and action-conditioned world simulation within a single model, enabling it to function as both an executable robot policy and a multi-embodiment visual world simulator. To scale embodied experience across heterogeneous data sources, we further develop a progressive embodied pretraining framework that unifies learning from egocentric human videos, handheld-gripper demonstrations, and heterogeneous robot trajectories under a shared World Action Modeling objective. Built upon 200K+ hours of interaction data, Riemann-1.0 progressively transfers large-scale embodied experience into executable robot manipulation capabilities. Riemann-1.0 achieves state-of-the-art performance across both simulation benchmarks and real-world manipulation tasks. It achieves success rates of 94.3% on RoboTwin2.0, 99.0% on LIBERO, and 62.6% on the long-horizon compositional benchmark RoboCasa-365, outperforming the previous best method by 8.4% On long-horizon real-world manipulation tasks, Riemann-1.0 achieves a Success Rate (SR) of 85.0% and a Progress Success Rate (PSR) of 94.4%, exceeding the strongest open-source baseline by 15% in SR. These results demonstrate that unified World Action Modeling together with progressive embodied pretraining effectively transforms large-scale embodied experience into generalizable robot manipulation capabilities.