DRHeC:基于RGB梯度的可微渲染手眼标定
DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients
浏览论文内容
中文总结 AI 辅助
提出基于RGB的可微渲染手眼标定框架DRHeC,融合颜色与掩码几何特征,提升精度与稳定性,在UR5e实验中抓取成功率88.9%,插入成功率57.4%,显著优于现有方法。
中文摘要 AI 辅助
准确的手眼标定对于精密操作至关重要。传统方法依赖于标记物,其精度取决于标记物的准确性和可观测性。相比之下,无标记方法(如基于学习的方法)使用深度神经网络直接从图像中提取关键点或特征,从而能够通过单张图像计算手眼变换,无需物理标记。近年来,基于可微渲染的手眼标定方法利用物理模型渲染二值掩码并与观测进行比较,从而在标定阶段无需基准标记即可实现手眼标定,并提供可解释的优化。尽管最先进的可微渲染方法达到了显著的精度,但使用二值掩码会导致内部轮廓细节的丢失,从而降低精度。此外,这些方法还可能遭受不稳定的优化和局部最小值问题。在本研究中,我们提出了一种新颖的基于RGB的可微渲染框架,通过结合颜色和掩码几何特征提供更丰富的几何和外观线索,从而提高标定精度和优化稳定性。此外,我们提出了一种掩码引导的图像到图像转换方法,以确保在转换过程中显式保留颜色和几何一致性。我们的方法通过仿真和真实世界实验进行了验证,结果表明其具有较高的精度和鲁棒性,并且明显优于现有的可微渲染方法。我们的方法在UR5e真实世界实验中实现了88.9%的抓取成功率和57.4%的插入成功率,分别比最先进的可微渲染手眼标定方法EasyHeC高出46.3和48.1个百分点。
英文摘要
Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations, enabling hand-eye calibration without fiducial markers in the calibration stage and providing interpretable optimization. While the state-of-the-art differentiable rendering methods achieve remarkable accuracy, the use of binary masks can result in the loss of internal profile details, reducing precision. Additionally, these methods can also suffer from unstable optimization and local minima. In this study, we propose a novel RGB-based differentiable rendering framework that provides richer geometric and appearance cues by incorporating color and mask geometric features, thereby improving calibration accuracy and optimization stability. Additionally, we propose a mask-guided image-to-image translation method to ensure explicit preservation of color and geometric consistency throughout the translation. Our approach is validated through both simulation and real-world experiments, with results demonstrating strong accuracy and robustness and clear improvements over existing differentiable rendering methods. Our method achieves a grasping success rate of 88.9% and insertion success rate of 57.4% on the UR5e real-world experiment, outperforming the state-of-the-art differentiable rendering hand-eye calibration method EasyHeC by 46.3 and 48.1 percentage points, respectively.
发表机构
- The University of Tokyo(东京大学)
- DENSO WAVE INCORPORATED(电装波株式会社)
机构由 AI 辅助整理,请以论文原文为准。