arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11270cs.ROcs.AI

迈向可预测、对齐且可扩展的机器人学习

Towards Predictive, Aligned, and Scalable Robot Learning

  • Astribot Team(Astribot团队)

机构由 AI 辅助整理,请以论文原文为准。

Peijun Tang, Shangjin Xie, Baifu Huang, Binyan Sun, Haotian Yang, Kuncheng Luo, Weiqi Jin, Shilin Fang, Jianan Wang

AI总结:

研究旨在实现可预测、对齐且可扩展的机器人学习。核心方法是引入Lumo - 2潜在世界 - 动作模型,提出多阶段模态预对齐策略。主要贡献是该模型在实证研究中表现出色,优于基线,验证了结构化多模态对齐和预测推理对具身智能的重要性。

AI中文摘要:

学习的核心超越记忆,具备通过在可能性空间中导航来推理和解决新问题的能力。我们引入Lumo - 2,一种潜在世界 - 动作模型,通过在潜在空间中对世界动态进行推理来生成动作。所学习的潜在世界动态捕捉基于物理的视觉过渡,自然编码未来可能性并为跨模态对齐提供统一基础。我们的方法核心假设是动作生成质量受潜在空间几何结构支配。为解决标准基于重建的动作令牌化目标导致的偏差问题,我们提出多阶段模态预对齐策略。我们对潜在世界建模和模态对齐进行了系统实证研究,结果表明Lumo - 2始终优于强大的视觉 - 语言 - 动作(VLA)和世界 - 动作模型(WAM)基线,在具有挑战性的现实世界任务中取得进展,这表明结构化多模态对齐和预测推理是推进具身智能的基本原则。

英文摘要:

Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities. We introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding future possibilities and providing a unified substrate for cross-modal alignment. This formulation enables predictive reasoning akin to world modelling while remaining lightweight and focused on physical dynamics relevant to control. Central to our approach is the hypothesis that action generation quality is governed by the geometry of the latent space. We observe that standard reconstruction-based action tokenization objectives induce representations biased toward low-level signal fidelity, leading to misalignment between reconstruction quality and downstream control performance. To address this limitation, we propose a multi-stage modality pre-alignment strategy in which action representations are progressively aligned with latent world dynamics, vision, and language. This process enforces cross-modal consistency, promotes abstraction, and induces a structured latent space for predictive reasoning. We provide a systematic empirical study of latent world modelling and modality alignment, analyzing their roles in scaling laws and out-of-distribution generalization. Results show that Lumo-2 consistently outperforms strong vision-language-action (VLA) and world-action model (WAM) baselines, with gains on challenging real-world tasks requiring temporal reasoning, physical understanding, or high control complexity, including long-horizon and dexterous manipulation. These findings suggest that structured multimodal alignment and predictive reasoning are fundamental principles for advancing embodied intelligence.

↑