arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

G6D:面向机器人操作的几何无学习RGB-D 6D姿态求解器

G6D: Geometric Learning-Free RGB-D 6D Pose Solver for Robotic Manipulation

Yixuan Liang, William Chen, Yunan Wang, Jizhou Yan, Zhao Jin, Changling Liu, Chuxiong Hu

arXiv 2609.23566首次发表:更新:

发表机构

Tsinghua University; Sapient Intelligence(清华大学; Sapient Intelligence)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器人操作中6D姿态估计对GPU资源依赖高、可解释性差的问题,提出无学习、几何驱动的G6D求解器,通过模板匹配与轮廓深度一致性优化,无需预训练模型,在多个数据集及真实实验中表现优异。

AI 中文摘要

6D物体姿态估计是机器人操作和自动化的基础。近期的零样本方法显著提高了对未见物体的泛化能力,但大多数仍依赖大规模预训练模型,需要大量的GPU计算和内存资源。这些需求使得在感知、规划和控制共享有限计算资源的机器人平台上部署变得复杂,而学习到的中间表示对于任务特定适配提供的几何可解释性有限。为解决这些限制,我们提出了G6D,一种无学习、几何驱动的RGB-D 6D姿态求解器。给定RGB-D观测、物体实例掩码、相机内参和CAD模型,G6D通过基于模板的几何匹配生成姿态假设,并使用轮廓和深度一致性进行细化,形成纯几何驱动的姿态估计范式。该范式既不需要预训练视觉模型,也不需要针对特定目标的训练,并在整个姿态估计过程中保持可解释的几何表示。此外,可调整的假设数量提供了灵活的精度-计算权衡,而仅CPU配置支持在无GPU资源的情况下部署。在LineMOD和五个BOP19数据集上的实验展示了先进的性能。真实世界的抓取放置实验进一步证明了G6D在机器人操作中的适用性。完整项目可在该https URL公开获取。

英文摘要

6D object pose estimation is fundamental to robotic manipulation and automation. Recent zero-shot methods have significantly improved generalization to unseen objects, but most still rely on large-scale pretrained models with substantial GPU computation and memory demands. These requirements complicate deployment on robotic platforms where perception, planning, and control share limited computational resources, while learned intermediate representations offer limited geometric interpretability for task-specific adaptation. To address these limitations, we propose G6D, a learning-free, geometry-driven RGB-D 6D pose solver. Given an RGB-D observation, an object instance mask, camera intrinsics, and a CAD model, G6D generates pose hypotheses through template-based geometric matching and refines them using silhouette and depth consistency, forming a purely geometry-driven pose estimation paradigm. This paradigm requires neither pretrained visual models nor target-specific training and preserves interpretable geometric representations throughout pose estimation. Moreover, adjustable hypothesis counts provide flexible accuracy-computation trade-offs, while a CPU-only configuration supports deployment without GPU resources. Experiments on LineMOD and five BOP19 datasets demonstrate advanced performance. Real-world pick-and-place experiments further demonstrate G6D's applicability to robotic manipulation. The complete project is publicly available at https://ai4control.github.io/G6D-Project-Page .

Comments9 pages, 5 figures. Corresponding author: Chuxiong Hu. Project page: https://ai4control.github.io/G6D-Project-Page

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑