arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02990cs.RO

EmbodiedVAE:用于高效可控具身操作的解耦视频变分自编码器

EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation

Jiayi Luo, Hanxin Zhu, Chen Gao, Jiankun Wang, Cong Wang, Tianyu He, Jianxin Li, Zhibo Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有潜在扩散模型适配具身操作场景的缺陷,提出EmbodiedVAE视频变分自编码器,通过双编码器架构与最优传输一致性模块实现更优重建质量、压缩率及精确动作控制。

中文摘要 AI 辅助

潜在扩散模型(LDMs)近期在构建强大的具身操作世界模型方面显著推进了具身学习。然而,尽管现有LDMs性能卓越,它们主要依赖针对自然场景优化的变分自编码器(VAEs),未考虑具身操作场景的独特特性,导致潜在表示既不紧凑也不可控,阻碍了LDMs的高效训练与精确机器人控制。为解决该问题,我们提出EmbodiedVAE,一种为机器人操作世界模型量身打造的新型视频VAE,可提供紧凑且可控的潜在表示。具体而言,EmbodiedVAE采用双编码器-单解码器架构,搭配非对称时空压缩模块,自动将机械臂运动与背景环境解耦,实现整体紧凑性的同时提供显式具身潜在以支持细粒度动作控制。为进一步保留学习到的机器人运动潜在的时间一致性,我们引入基于最优传输的一致性模块,明确强化运动保真度与帧间连贯性。大量实验表明,所提出的EmbodiedVAE实现了出色的重建质量与高压缩率,同时在机器人操作场景中支持更精确的动作控制,相比最先进的视频VAEs平均实现2dB的峰值信噪比(PSNR)提升。

英文摘要

Latent diffusion models (LDMs) have recently significantly advanced embodied learning in constructing powerful embodied manipulation world models. However, despite the remarkable performance, existing LDMs predominantly rely on Variational Autoencoders (VAEs) optimized for natural scenes while failing to account for the unique characteristics of embodied manipulation scenarios, yielding latent representations that are neither compact nor controllable, thereby hindering efficient training of LDMs and precise robotic control. To solve this problem, we present EmbodiedVAE, a novel video VAE that provides compact yet controllable latent representations tailored for the robotic manipulation world models. Specifically, EmbodiedVAE adopts a dual-encoder, single-decoder architecture with an asymmetric spatio-temporal compression module, which automatically disentangles the robot arm's motion from background environment, resulting in overall compactness while providing explicit embodied latent to support fine-grained action control. To further preserve the temporal consistency of learned robotic motion latent, we introduce an optimal-transport-based consistency module that explicitly enforces motion fidelity and inter-frame coherence. Extensive experiments demonstrate that our proposed EmbodiedVAE achieves superior reconstruction quality with high compression rate, while enabling more precise action control in robotic manipulation scenarios with an average of 2dB PSNR improvement over state-of-the-art video VAEs.

发表机构

  • Beihang University(北京航空航天大学)
  • Zhongguancun Academy(中关村学院)
  • University of Science and Technology of China(中国科学技术大学)
  • National University of Singapore(新加坡国立大学)
  • CASIA, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所(CASIA))
  • Microsoft Research Asia(微软亚洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑