arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WLA$^3$:面向语义、动力学和运动学的世界潜在动作建模

WLA$^3$: World Latent Action Modeling for Semantics, Dynamics, and Kinematics

Peidong Liu, Zhiyuan Xiang, Mingyang Li, Wenhao Li, Jiale Zhang, Jiahao Sun, Jiawei Li

arXiv 2609.15870首次发表:更新:

发表机构

Joy Future Academy, JD Group(京东探索研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

WLA$^3$通过世界潜在动作模型统一学习语义、动力学和运动学表示,利用人类视频和机器人轨迹,在LARYBench和真实机器人任务中显著提升性能。

AI 中文摘要

使用异构数据扩展通用策略模型受到缺乏统一、低噪声动作监督的限制。人类第一人称视频数量丰富,但只有一小部分带有高质量的手部动作标签。观察到的世界状态转换提供了跨数据源的常见动作相关监督来源。我们提出了WLA$^3$(面向语义、动力学和运动学的世界潜在动作建模),这是一个统一的通用策略模型框架,围绕世界潜在动作模型(WLAM)学习到的表示构建。WLAM首先学习多模态世界状态在局部时间间隔内的变化方式,将同步的相机视图和可用的具身状态变化编码为紧凑的局部潜在动作和更丰富的转换特征。从部分模态进行重建以及重叠窗口之间的一致性促进了鲁棒的转换表示。WLA$^3$在语义、动力学和运动学中重用这些表示:局部潜在动作支持对动作敏感的物理动力学建模,段级特征通过语义潜在聚合(SLA)直接监督视觉语言模型(VLM),动作专家联合预测潜在动作以及特定于具身的机器人控制。人类视频提供可扩展的转换监督,而机器人轨迹将共享表示锚定在可执行的原生控制中。在LARYBench上,最终的32维潜在动作达到了67.89%的平均分类准确率。WLA$^3$在六个真实机器人任务中实现了81.9%的平均成功率,而$\pi_{0.5}$为66.2%。性能随着通用策略模型中期训练数据的扩展而提高,人类视频支持人机迁移。项目页面可在本https URL找到。

英文摘要

Scaling generalist policy models with heterogeneous data is limited by the lack of unified, low-noise action supervision. Human egocentric videos are abundant, but only a small fraction comes with high-quality hand-action labels. Observed world transitions offer a common source of action-related supervision across data sources. We introduce WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM). WLAM first learns how multimodal world states change over a local interval, encoding synchronized camera views and available embodiment-state changes into a compact local latent action and a richer transition feature. Reconstruction from partial modalities and consistency across overlapping windows encourage robust transition representations. WLA$^3$ reuses them across semantics, dynamics, and kinematics: local latent actions support action-sensitive physical-dynamics modeling, segment-level features directly supervise the VLM through a Semantic Latent Aggregate (SLA), and an action expert jointly predicts latent actions together with embodiment-specific robot controls. Human videos provide scalable transition supervision, while robot trajectories ground the shared representation in executable native controls. On LARYBench, the final 32D latent action reaches 67.89\% average classification accuracy. WLA$^3$ achieves 81.9% average success across six real-robot tasks versus 66.2% for $π_{0.5}$. Performance improves as generalist policy model mid-training data scales, and human videos support human-to-robot transfer. Project page can be found at https://wla-3.github.io/.

CommentsProject page can be found at https://wla-3.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑