arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.11832cs.RO

基于多视图潜在先验的学习动作流形用于机器人操作

Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation

Junjin Xiao, Dongyang Li, Yandan Yang, Shuang Zeng, Tong Lin, Xinyuan Chang, Feng Xiong, Mu Xu, Xing Wei, Zhiheng Ma, Qing Zhang, Wei-Shi Zheng

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出利用多视图扩散模型合成潜在新视角,并引入几何引导门控变换器以提升机器人视觉-语言-动作模型的空间感知与操作性能,通过动作流形学习提高动作学习效率,实验证明优于现有方法。

中文摘要 AI 辅助

本文针对视觉-语言-动作(VLA)模型中的空间感知和操作挑战,通过预训练的多视图扩散模型合成潜在新视角,并提出几何引导门控变换器(G3T),在3D几何引导下对齐多视角特征并自适应过滤遮挡噪声。为提高动作学习效率,引入动作流形学习(AML),直接在有效动作流形上预测动作,绕过对无结构目标如噪声或速度的低效回归。在LIBERO、RoboTwin 2.0和真实机器人任务上的实验表明,本文方法在成功率和鲁棒性上优于现有最佳方法。项目页面:https://junjxiao.github.io/Multi-view-VLA.github.io/.

英文摘要

This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views and propose a Geometry-Guided Gated Transformer (G3T) that aligns multi-view features under 3D geometric guidance while adaptively filtering occlusion noise. To improve action learning efficiency, we introduce Action Manifold Learning (AML), which directly predicts actions on the valid action manifold, bypassing inefficient regression of unstructured targets like noise or velocity. Experiments on LIBERO, RoboTwin 2.0, and real-robot tasks show our method achieves superior success rate and robustness over SOTA baselines. Project page: https://junjxiao.github.io/Multi-view-VLA.github.io/.

发表机构

  • Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China(机器智能与先进计算国家重点实验室,教育部,中国)
  • AMap, Alibaba Group(阿里的AMap)
  • Xi’an Jiaotong University(西安交通大学)
  • Shenzhen University of Advanced Technology(深圳先进技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑