arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LeapBot-WA:通过预测性潜在对齐实现的世界锚定动作模型

LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

Pei Liu, Nan Zheng, Lang Zhang, Daojie Peng, Yanan Zhang, Feilong Kong, Mingyue Feng, Jiachao Liu, Yaonong Wang, Qifeng Chen, Jun Ma

arXiv 2607.23969首次发表:更新:

AI 中文总结

研究针对世界动作模型依赖像素级视频生成的瓶颈,提出LeapBot-WA,通过预测性潜在对齐建立新范式,引入ISAE弥合模态差距,设计MoT架构,在多数据集上表现出色,实现高效强大的潜在中心范式及零样本鲁棒性和现实世界迁移。

AI 中文摘要

世界动作模型(WAMs)已成为具身智能的强大范式,但对像素级视频生成的普遍依赖造成了根本瓶颈。强迫模型重建与任务无关的视觉细节会消耗表征能力,使策略易受视觉干扰。本文提出LeapBot-WA,通过将联合嵌入预测架构(JEPA)作为世界锚定来建立一种新的预测性潜在范式。它将世界建模核心转向预测语义对齐,在潜在基础空间中直接提取抽象物理动力学。为弥合非高斯预测特征与扩散先验之间的模态差距,引入各向同性语义自动编码器(ISAE)。还设计了非对称混合变压器(MoT)架构。训练时,锚定扩散变压器指导动作扩散变压器;推理时修剪重动力学分支。LeapBot-WA在LIBERO上的预测模型中取得了领先性能,在RoboTwin 2.0上与顶级生成式WAMs相当,无需大规模轨迹预训练,还展示了对未知环境的卓越零样本鲁棒性和成功的现实世界迁移,为可扩展机器人控制建立了高效且强大的以潜在为中心的范式。

英文摘要

World Action Models (WAMs) have emerged as a powerful paradigm for embodied intelligence, yet the prevailing reliance on pixel-level video generation creates a fundamental bottleneck. Forcing models to reconstruct task-irrelevant visual details dissipates representational capacity and renders policies vulnerable to visual distractors. In this paper, we propose LeapBot-WA, which establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor. Departing from the traditional reliance on visual synthesis, LeapBot-WA shifts the core of world modeling to Predictive Semantic Alignment, extracting abstract physical dynamics directly within a latent foundation space. To bridge the modality gap between non-Gaussian predictive features and diffusion priors, we introduce the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift. Furthermore, we design an Asymmetric Mixture-of-Transformers (MoT) architecture. During training, an Anchor Diffusion Transformer acts as a privileged dynamics expert to guide the Action Diffusion Transformer; at inference, this heavy dynamics branch is pruned, enabling zero-overhead execution. LeapBot-WA achieves state-of-the-art performance among predictive models on LIBERO and matches top-tier generative WAMs on RoboTwin 2.0 without requiring large-scale trajectory pre-training. It further demonstrates superior zero-shot robustness to unseen environments and successful real-world transfer, establishing a highly efficient and robust latent-centric paradigm for scalable robotic control. Code: https://github.com/LeapWM/leapbot-wa.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑