ARGUS:利用大规模3D视觉模型在视角变化下对齐机器人场景几何结构
ARGUS: Aligning Robot Scene Geometry Under Shifting Views with Large 3D Vision Models
浏览论文内容
中文总结 AI 辅助
本研究提出ARGUS,一种利用大规模3D视觉模型对齐机器人场景几何结构的观测预处理流水线,可提升机器人操纵策略在视角多样化场景下的性能与学习效率。
中文摘要 AI 辅助
大规模视觉运动策略在各类机器人操纵任务中展现出了出色性能。然而,尽管取得了这些成功,操纵策略常将场景几何结构与对应视角绑定,学习的是物体在图像中的位置而非其在任务空间中的位置。这种绑定固有地限制了对应策略从视角多样化数据集(如DROID、BridgeV2)中学习的能力,也限制了其泛化到训练数据所捕捉视角之外的能力。本研究中,我们提出ARGUS,一种观测预处理流水线,该流水线利用大规模3D视觉模型将任意相机视角的图像观测对齐到一个标准视角,再将其传递给下游视觉运动策略。在具有不同视角多样性水平的训练数据集上开展的实验,从固定多相机配置到高度多样化的相机布置,均表明我们的方法在有限视角和视角多样化训练场景下始终优于现有方法。在效率对比中,ARGUS展现出从视角多样化数据中学习的能力,通过利用简化的观测空间,其收敛到高成功率的速度比先前方法快4至6倍。总体而言,我们的发现表明,利用大规模3D视觉模型可减轻视觉运动策略的学习负担,使其能从大规模、视角多样化的机器人数据中更高效地学习。
英文摘要
Large-scale visuomotor policies have demonstrated impressive performance across a wide range of robot manipulation tasks. However, despite this success, manipulation polices often entangle scene geometry with the corresponding viewpoint, learning where objects lie in an image rather than where it lies in the task space. This entanglement inherently limits the corresponding policy's ability to learn from viewpoint-diverse datasets (ex. DROID, BridgeV2) and generalize beyond the viewpoints captured in their training data. In this work, we present ARGUS, an observation pre-processing pipeline that uses large-scale 3D vision models to align image observations from arbitrary camera viewpoints into a canonical viewpoint before passing it to downstream visuomotor policies. Experiments across training datasets with varying levels of viewpoint diversity, from fixed multi-view camera configurations to highly varied camera placements, show that our method consistently outperforms prior approaches across both limited-view and view-diverse training regimes. In efficiency comparisons, ARGUS demonstrates an ability to learn from view-diverse data, converging to high success rates 4-6x faster than previous methods by leveraging a simplified observation space. Overall, our findings show that leveraging large-scale 3D vision models reduces the learning burden on visuomotor policies, enabling more efficient learning from large-scale, viewpoint-diverse robot datasets.
发表机构
- University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。