MobileVISTA:移动操作中姿态泛化的生成式数据增强
MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation
浏览论文内容
中文总结 AI 辅助
MobileVISTA通过联合增强视觉观测和重新定向动作,生成姿态扰动的训练数据,提升移动操作策略对部署时姿态偏差的鲁棒性,无需额外示范或生成模型。
中文摘要 AI 辅助
人形机器人等移动操作器正越来越多地被部署在动态、非结构化环境中执行灵巧操作任务。然而,从单一机器人姿态采集的示范数据训练出的端到端操作策略是脆弱的:在部署时,即使机器人姿态出现厘米级的偏差,也会使以自我为中心的观测和末端执行器轨迹偏离训练分布,导致性能急剧下降。我们提出了MobileVISTA,一个数据生成框架,通过联合(1)增强以自我为中心的视觉观测和(2)重新定向动作以补偿基座姿态变化,将规范姿态下采集的示范转换为多样化的、姿态扰动的训练数据。与先前假设相机刚性安装在非驱动链上或假设复杂的铰接机器人几何体大部分在画面之外的方法不同,MobileVISTA旨在兼容以自我为中心的平台(如人形机器人),在这些平台上,相机既受机器人运动链的影响,又必须在机器人运动时观察其运动链。我们在涵盖人形和双臂形态的模拟任务中,以及在真实的Galaxea R1 Pro上研究了MobileVISTA。我们发现,在MobileVISTA增强数据上训练的策略,对测试时遇到的先前分布外姿态表现出更强的鲁棒性,而无需额外采集示范或使用训练好的生成模型。此外,我们发现MobileVISTA的收益在测试的人形机器人上最大,因为在这些机器人上,相机搭载在驱动链上,且机器人占据了画面的大部分。更多视频和附录可在我们的网站上找到:此https URL
英文摘要
Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot pose at deployment can drive ego-centric observations and end-effector trajectories out of the training distribution, leading to sharp drops in performance. We introduce MobileVISTA, a data generation framework that transforms demonstrations captured at canonical poses into diverse, pose-perturbed training data by jointly (1) augmenting egocentric visual observations and (2) retargeting actions to compensate for base pose changes. Unlike prior methods, which assume a camera rigidly mounted off the actuated chain or non-trivial articulated robot geometry largely out of frame, MobileVISTA targets compatibility with egocentric platforms (e.g., humanoids) where the camera is both influenced by and must observe the robot's kinematic chain as it moves. We study MobileVISTA in simulated tasks spanning humanoid and bimanual embodiments, and on a real Galaxea R1 Pro. We find policies trained on MobileVISTA-augmented data demonstrate improved robustness to previously out-of-distribution poses encountered at test time, without additional demonstration collection or a trained generative model. Additionally, we find MobileVISTA's benefit is largest on tested humanoids, where the camera rides the actuated chain and the robot fills much of the frame. Additional videos and appendix can be found on our website: https://mobilevista.github.io
发表机构
- Stanford University(斯坦福大学)
- Toyota Research Institute(丰田研究所)
机构由 AI 辅助整理,请以论文原文为准。