arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38443cs.ROcs.CV

BIND:将3D机器人动作绑定到2D图像特征

BIND: Binding 3D Robot Actions to 2D Image Features

Cameron Smith, Arsh Tangri, Vitor Guizilini, Yue Wang, Zubair Irshad, Sergey Zakharov

首次发表
浏览论文内容

中文总结 AI 辅助

BIND通过相机几何将3D机器人动作与2D图像特征绑定,替代全局特征回归,实现高数据效率(5次演示即可)和对分布外物体位置及相机视角的强鲁棒性。

中文摘要 AI 辅助

我们提出了BIND,一种用于视觉运动机器人策略的新动作表示方法,它将3D机器人动作与其对应的2D图像特征绑定,从而带来显著的数据效率提升以及对分布外物体位置和相机视角的鲁棒性。当前机器人策略的动作头通常被构建为从预训练视觉编码器产生的单一全局特征向量进行MLP回归。这种全局公式要求策略网络仅从演示中自行发现目标机器人动作与其投影到的图像特征之间的关系。结果是,尽管现代图像特征具有语义描述性、空间鲁棒性,甚至多视角一致性,但基于这些特征构建的策略对相机视角和物体放置的细微变化非常脆弱,而且数据效率出奇地低。BIND通过相机几何而非学习来提供动作-特征关系,从而弥合了这一差距:它离散化候选末端执行器位置的体积,将每个候选位置附着到其在每个相机视图投影处的预训练特征上,并通过评分每个候选位置及其图像绑定特征的组合来选择动作。在真实机器人上,我们研究了数据效率、对未见物体位置和相机视角的分布外鲁棒性,以及一般的长时程任务执行和灵巧性。我们发现BIND具有高度数据效率和鲁棒性:在仅5次演示的任务上就能达到近乎完美的成功率,并且在陡峭的相机视角偏移和保留物体位置下优雅地退化,而坐标回归基线在这些情况下完全失败。

英文摘要

We introduce BIND, a new action representation for visuomotor robot policies that binds 3D robot actions to their corresponding 2D image features, yielding strong data efficiency gains and robustness to out-of-distribution object positions and camera viewpoints. The action heads of current robot policies are typically formulated as an MLP regression from a single global feature vector produced by a pre-trained vision encoder. This global formulation requires the policy network to discover, from demonstrations alone, the relationship between target robot actions and the image features they project onto. The consequence is that although modern image features are semantically descriptive, spatially robust, and even multiview-consistent, the policies built on them are brittle to subtle changes in camera viewpoint and object placement--and surprisingly data-inefficient. BIND closes this gap by supplying the action-feature relationship through camera geometry rather than learning: it discretizes a volume of candidate end effector positions, attaches each candidate to the pre-trained features at its projection in each camera view, and selects actions by scoring each candidate's position and image-bound feature combination. On a real robot, we study data efficiency and out-of-distribution robustness to unseen object positions and camera viewpoints, as well as general long-horizon task execution and dexterity. We find BIND to be highly data-efficient and robust: it achieves near-perfect success on tasks with as few as 5 demonstrations, and degrades gracefully under steep camera-viewpoint shifts and held-out object positions where coordinate-regression baselines completely fail.

发表机构

  • University of Southern California(南加州大学)
  • Toyota Research Institute(丰田研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑