arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15509cs.RO

StereoPatch:用于机器人操作中空间感知的补丁对齐RGB-深度融合

StereoPatch: Patch-Aligned RGB-Depth Fusion for Spatial Perception in Robot Manipulation

  • Australian Centre for Robotics, The University of Sydney(悉尼大学澳大利亚机器人中心)

机构由 AI 辅助整理,请以论文原文为准。

Yanan Zhou, Zhaoyan Qian, James Zhao, Weiming Zhi

AI总结:

StereoPatch通过补丁对齐的RGB-深度融合,将度量几何直接绑定到RGB特征,提升机器人操作中的空间感知与闭环成功率。

AI中文摘要:

机器人模仿学习的最新进展产生了直接从视觉观测预测动作的视觉运动策略。然而,当目标位置、物体高度或接触几何形状变化时,视觉上相似的场景可能需要不同的动作。预训练的RGB特征可能将这些几何上不同的状态映射到相似的政策输入,而简单地添加深度则要求策略从用于学习控制的相同有限演示中学习RGB-深度对应关系。我们引入了StereoPatch,一种补丁对齐的RGB-深度表示,它将注册的度量几何直接绑定到用于动作预测的RGB补丁上。在共享的2D补丁网格上,不对称交叉注意力在动作解码之前将深度信息融入相应的RGB特征中。由此产生的StereoPatch令牌提供了一种几何感知的视觉表示,可以在不改变其底层学习目标的情况下调节通用视觉运动策略。在六个真实机器人任务中,StereoPatch实现了比仅外观、仅几何、原始RGB-D和晚期融合基线更高的闭环成功率。额外的实验在三个模拟套件中评估了跨视觉运动策略架构的兼容性、空间泛化和操作极限。结果表明,解决控制相关的几何模糊性受益于将深度直接与用于动作预测的视觉特征对齐,而不是将其作为独立模态提供。项目页面:此https URL。

英文摘要:

Recent advances in robot imitation learning have produced visuomotor policies that predict actions directly from visual observations. Yet visually similar scenes can require different actions as target position, object height, or contact geometry changes. Pretrained RGB features may map these geometrically distinct states to similar policy inputs, while simply adding depth requires the policy to learn RGB-depth correspondence from the same limited demonstrations used to learn control. We introduce StereoPatch, a patch-aligned RGB-depth representation that binds registered metric geometry directly to the RGB patches used for action prediction. On a shared 2-D patch grid, asymmetric cross-attention incorporates depth information into the corresponding RGB features before action decoding. The resulting StereoPatch Tokens provide a geometry-aware visual representation that can condition general visuomotor policies without changing their underlying learning objectives. Across six real-robot tasks, StereoPatch achieves higher closed-loop success than appearance-only, geometry-only, raw RGB-D, and late-fusion baselines. Additional experiments across three simulation suites evaluate compatibility across visuomotor policy architectures, spatial generalization, and operating limits. Results suggest that resolving control-relevant geometric ambiguity benefits from aligning depth directly with the visual features used for action prediction, rather than supplying it as an independent modality. Project page: https://aus.bot/research/stereopatch/.

补充信息

↑