arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无先验的多物体实例相对6D姿态估计

Prior-free relative 6D pose estimation of multiple object instances

Behdad Khodabandehloo, Andrea Caraffa, Davide Boscaini, Fabio Poiesi

arXiv 2609.08949首次发表:更新:

发表机构

Fondazione Bruno Kessler; University of Trento(布鲁诺·凯斯勒基金会; 特伦托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无先验相对6D姿态估计新设定,无需CAD模型或参考图像,通过多模态基础特征与循环一致性细化对应关系,估计未知物体多实例间相对姿态,新方法PROSE在基准PRENCH上超越现有方法。

AI 中文摘要

物体6D姿态估计的公式化方法已逐步减少对物体特定先验的依赖,从显式3D模型发展到多视角物体捕获,再到单一参考图像。我们将这一进展推向极致,引入了无先验的相对6D姿态估计,该设定摒弃了已知场景中待估计姿态的物体是哪一类的假设。这一新设定旨在估计同一图像中未知物体的多个实例之间的相对姿态,无需CAD模型、模板或参考图像。我们通过提出一种新方法(PROSE)来解决这一问题,该方法利用多模态基础特征在物体实例之间找到粗略对应关系,因此无需训练。我们通过跨实例元组施加循环一致性来细化这些对应关系,并利用由此产生的全局一致对应关系来估计任意一对实例之间的相对6D姿态。为了进行系统性评估,我们设计了一个新基准(PRENCH),该基准基于三个多实例BOP数据集构建,并丰富了任务特定的元数据。PROSE在适应所提设定的最先进单图像方法所获得的基线上持续取得更优表现,同时既不需要任务特定的监督,也不需要额外的学习组件。项目网站:此HTTPS URL。

英文摘要

Object 6D pose estimation formulations have progressively reduced reliance on object-specific priors, evolving from explicit 3D models to multi-view object captures to single reference images. We take this progression to its extreme by introducing prior-free relative 6D pose estimation, which lifts the assumption of knowing which object is to be posed within the scene. This novel setting aims to estimate the relative poses of multiple instances of an unknown object within the same image, without requiring CAD models, templates, or reference images. We solve this by formulating a novel method (PROSE) that finds coarse correspondences between object instances using multimodal foundation features, thus requiring no training. We refine these correspondences by imposing cycle consistency across tuples of instances, and leverage the resulting globally consistent correspondences to estimate the relative 6D pose between any pair of instances. To enable systematic evaluation, we design a novel benchmark (PRENCH) built from three multi-instance BOP datasets and enriched with task-specific metadata. PROSE consistently outperforms baselines obtained by adapting state-of-the-art single-image methods to the proposed setting, while requiring neither task-specific supervision nor additional learned components. Project website: https://tev-fbk.github.io/PROSE/

CommentsTechnical report. 12 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑