VGM-VS:重新思考用于高精度视觉伺服的可视几何模型
VGM-VS: Rethinking Visual Geometry Model for High-Precision Visual Servoing
浏览论文内容
中文总结 AI 辅助
VGM-VS利用预训练可视几何模型进行闭环位姿视觉伺服,通过场景特定度量自适应解决尺度模糊,实现亚毫米精度和90-100%成功率,适用于高公差装配任务。
中文摘要 AI 辅助
我们提出了VGM-VS,一种基于预训练前馈可视几何模型的视觉伺服方法。给定当前视图和在目标配置下拍摄的参考图像,我们使用可视几何模型估计相对相机位姿,并将其迭代地作为闭环基于位姿的视觉伺服(PBVS)方案的位姿增量。从大规模预训练中获得的几何感知表示在目标被遮挡、纹理较弱或仅覆盖图像小部分时,仍能保持该估计的可靠性。然而,这些模型固有的尺度模糊性使得预测的平移仅定义到未知尺度,而位姿增量必须具有度量性才能用于机器人控制。我们通过场景特定的度量自适应来弥合这一差距:机器人从目标位姿开始,沿预定义运动自主记录图像-位姿对,我们在此数据上微调相机头部,联合学习手眼变换,从而无需专门的标定过程。我们在三个具有苛刻公差的真实世界装配任务上评估了我们的方法:USB-C电缆拾取、电缆插入和内存条插入。VGM-VS以30Hz实时运行,在电缆任务上收敛到亚毫米级终端精度,并且在伺服过程中目标移动时达到90-100%的成功率。在初始位移高达30cm且目标物体50%被遮挡的情况下,它在所有试验中均收敛,优于所比较的视觉伺服基线。
英文摘要
We present VGM-VS, a visual servoing method built on a pretrained feed-forward visual geometry model. Given the current view and a reference image captured at the target configuration, we estimate the relative camera pose with a visual geometry model and apply it iteratively as the pose increment of a closed-loop pose-based visual servoing (PBVS) scheme. The geometry-aware representation acquired from large-scale pretraining keeps this estimate reliable when the target is occluded, weakly textured, or covers only a small part of the image. However, the scale ambiguity inherent to these models leaves the predicted translation defined up to an unknown scale, while the pose increment must be metric for robot control. We close this gap with a scene-specific metric adaptation: the robot autonomously records image--pose pairs along a predefined motion starting from the target pose, and we fine-tune the camera head on these data, jointly learning the hand--eye transform and thus removing the need for a dedicated calibration process. We evaluate our method on three real-world assembly tasks with demanding tolerances: USB-C cable picking, cable insertion, and RAM insertion. Running in real time at 30Hz, VGM-VS converges to submillimeter terminal accuracy on the cable tasks, and reaches success rates of 90--100\% when the target is moved during servoing. It converges in all trials under initial displacements of up to 30cm from the reference pose and with 50\% of the target object occluded, outperforming the compared visual servoing baselines.
发表机构
- Agile Robots SE(敏捷机器人公司)
- Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。