发表机构
University of Science and Technology of China; Beijing Institute of Control Engineering(中国科学技术大学; 北京控制工程研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PriorPose提出参考引导的联合变形与对齐框架,显式利用类别先验,在共享特征空间联合预测NOCS场和规范变形,以提升类别级物体姿态估计的精度与鲁棒性。
AI 中文摘要
类别级物体姿态估计旨在无需实例特定CAD模型的情况下,为未见实例恢复相似变换$(R,t,s)$。大多数有竞争力的方法基于对应关系:无先验变体直接从局部观测回归规范(NOCS)坐标,并在权重中隐式记忆规范框架,这使参数与类别典型方向绑定,并在分布偏移下损害泛化能力;有先验变体引入类别先验,但通常采用串行的先变形后对齐流程,其中欠约束的规范补全会破坏对应关系并导致姿态中的错误级联。我们提出PriorPose,一种参考引导的对应关系框架,保持类别先验显式,并在共享特征空间中联合求解规范化和对齐。参考引导的种子变换器将部分观测和类别先验作为令牌集嵌入,并通过几何感知种子融合它们,网络从中联合预测可见点的逐点NOCS场和先验的规范变形以重建完整规范实例,同时深度姿态头从诱导的对应关系回归$(R,t,s)$。一个两部分形状一致性目标,包含规范空间和相机空间一致性损失,耦合对应关系、变形和姿态,减少对记忆的规范方向的依赖,并避免先变形后对齐的错误级联。在标准和更大类别基准上的实验表明,PriorPose在大多数评估指标上创下新的最先进结果,尤其在严格姿态阈值下,同时在宽松姿态和IoU指标上保持竞争力,并在形状变化和域偏移下展现出改进的鲁棒性。
英文摘要
Category-level object pose estimation seeks to recover a similarity transform $(R,t,s)$ for unseen instances without instance-specific CAD models. Most competitive methods are correspondence-based: prior-free variants regress canonical (NOCS) coordinates directly from local observations and implicitly memorize the canonical frame in the weights, which ties the parameters to category-typical orientations and hurts generalization under distribution shift; prior-based variants introduce a category prior but typically follow a serial deform-then-align pipeline, where underconstrained canonical completion can corrupt correspondences and induce error cascades in pose. We propose PriorPose, a reference-guided correspondence framework that keeps the category prior explicit and solves canonicalization and alignment jointly in a shared feature space. A reference-guided seeded transformer embeds the partial observation and the category prior as token sets and fuses them via geometry-aware seeds, from which the network jointly predicts a per-point NOCS field for visible points and a canonical deformation of the prior that reconstructs a full canonical instance, while a deep pose head regresses $(R,t,s)$ from the induced correspondences. A two-part shape consistency objective, with canonical-space and camera-space consistency losses, couples correspondence, deformation, and pose, reducing reliance on memorized canonical orientations and avoiding deform-then-align error cascades. Experiments on standard and larger-category benchmarks demonstrate that PriorPose sets new state-of-the-art results on most evaluated metrics, especially under strict pose thresholds, while remaining competitive on relaxed pose and IoU metrics and showing improved robustness under shape variation and domain shift.
CommentsAccepted to ECCV 2026