arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01758cs.CV

GenCOPE:面向机器人抓取的Syn2Real广义类别级物体姿态估计

GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

Jian Liu, Wei Sun, Zhenqi Dai, Hui Yang, Jian Xiao, Nicu Sebe, Na Zhao

AI总结:

提出GenCOPE,仅用合成数据训练实现类别级物体姿态估计的域泛化,通过2D/3D语义一致性与交叉一致性学习,以轻量高效架构在真实场景中取得优越性能。

AI中文摘要:

类别级物体姿态估计(COPE)能够泛化到同类未知物体,已成为机器人三维场景理解的核心技术。然而,现有的COPE方法仍需要对新物体类别进行费力的真实世界训练数据重新采集,这限制了其在实际应用中的可扩展性。本文旨在实现合成到真实(Syn2Real)的广义COPE,即模型仅基于渲染的合成数据训练,并直接泛化到真实世界部署。核心挑战在于合成数据与真实数据之间存在显著的域差距,尤其是在纹理外观方面。为解决此问题,我们旨在通过学习域不变表示来增强域泛化,该表示捕获同一类别内物体的语义共性。我们引入了2D和3D语义一致性约束,以降低特征编码器对域特定特征的敏感性。此外,我们提出了一种端到端的姿态回归框架,该框架执行2D-3D交叉一致性学习,利用密集的跨模态融合进一步细化姿态估计。由于简单性和有效性对于真实世界的机器人部署至关重要,我们的模型仅基于全局特征运行,从而产生高度轻量且高效的架构。在REAL275和Wild6D基准以及真实世界机器人操作场景上的大量实验表明,我们的范式具有优越的Syn2Real泛化性能。代码和演示已在此https URL发布。

英文摘要:

Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scalability in practical applications. This paper aims to achieve synthetic-to-real (Syn2Real) generalized COPE, where a model is trained solely on rendered synthetic data and directly generalized to real-world deployments. The central challenge lies in the significant domain gap between synthetic and real-world data, particularly in texture appearance. To address this, we aim to enhance domain generalization by learning domain-invariant representations that capture semantic commonalities among objects within the same category. We introduce 2D and 3D semantic consistency constraints to reduce the sensitivity of feature encoders to domain-specific features. In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Since simplicity and effectiveness are essential for real-world robotic deployment, our model operates exclusively on global features, yielding a highly lightweight and efficient architecture. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm. Code and demos are released at https://paperreview99.github.io/GenCOPE/.

补充信息

↑