arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06726cs.CV

回归特征:基于稠密局部特征的零样本6DoF姿态估计

Back to the Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features

Ali Rafiaei, Michael Greenspan

首次发表
浏览论文内容

中文总结 AI 辅助

B2TFPose利用冻结的DINOv3提取稠密局部特征,通过测地线NMS、渲染引导重新对应和多掩膜选择,实现无需训练的零样本6DoF姿态估计,在BOP基准上达到最先进性能。

中文摘要 AI 辅助

我们提出了B2TFPose,一种无需训练、零样本的6DoF姿态估计方法,用于从RGB图像中估计未见物体的姿态。在姿态估计流程中,B2TFPose仅使用一个冻结的DINOv3视觉Transformer作为唯一的预训练组件,提取稠密的块级特征,这些特征无需任何任务特定的微调即可跨越合成到真实的域差距,通过大规模自监督基础模型的视角重新审视经典的局部特征匹配范式。三项贡献推动了无需训练方法的最新技术水平。一种测地线非极大值抑制策略检索视角多样的模板集,用于从粗到细的对应匹配。渲染引导的重新对应(RRC)在估计姿态处合成物体特定视图,并重新建立稠密的2D-3D对应关系,以在不增加学习参数的情况下锐化初始估计。一种多掩膜假设选择策略联合评分竞争的候选分割掩膜,以解决分割模糊性。在BOP基准的七个核心数据集上,B2TFPose在无细化时达到40.7的平均AR,在有细化时达到56.4,在无需训练的RGB方法中确立了最先进的性能,并以具有竞争力的推理速度超越了包括GigaPose和GenFlow在内的训练方法。

英文摘要

We present B2TFPose, a training-free zero-shot method for 6DoF pose estimation of unseen objects from RGB images. Using a single frozen DINOv3 vision transformer as its only pretrained component within the pose estimation pipeline, B2TFPose extracts dense patch-level features that generalize across the synthetic-to-real domain gap without any task-specific fine-tuning, revisiting the classical local feature matching paradigm through the lens of large-scale self-supervised foundation models. Three contributions advance the training-free state of the art. A geodesic non-maximum suppression strategy retrieves a viewpoint-diverse template set for coarse-to-fine correspondence matching. Render-guided Re-Correspondence (RRC) synthesizes object-specific views at the estimated pose and re-establishes dense 2D-3D correspondences to sharpen the initial estimate without additional learned parameters. A multi-mask hypothesis selection strategy jointly scores competing segmentation candidates to resolve segmentation ambiguity. On the seven core datasets of the BOP Benchmark, B2TFPose achieves 40.7 mean AR without refinement and 56.4 with refinement, establishing state-of-the-art performance among training-free RGB methods and outperforming trained counterparts including GigaPose and GenFlow, at competitive inference speed.

发表机构

  • Queen’s University(女王大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑