arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19872cs.CVcs.GR

PART:利用Transformer学习3D零件装配与检索

PART: Learning 3D Part Assembly and Retrieval with Transformers

Ruchao Bao, Wenzheng Wu, Chucheng Xiang, Zhongyuan Liu, Yuan Liu, Jinxin Dong, Ligang Liu, Ziqi Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出统一Transformer框架PART,将3D零件检索与装配视为集合预测,解决搜索空间爆炸、变长输出和6自由度姿态估计,通过联合训练和分割增强优化,在8万+形状数据集上验证泛化性。

中文摘要 AI 辅助

3D装配是现代制造和数字内容创作的基础。在本文中,我们提出了PART,一个统一的基于Transformer的3D零件检索与装配框架:给定一个目标形状和一个零件库,PART自动选择合适的零件并预测其6自由度姿态以重建目标。尽管先前的工作在装配预定义零件集方面取得了显著进展,但这种更实用的基于检索的设置仍未得到充分探索。该任务面临三个关键挑战:(i)组合爆炸式的搜索空间,随库大小呈指数增长;(ii)变长输出,因为不同目标需要不同数量的零件;(iii)用于零件装配的连续6自由度姿态估计。为解决这些问题,我们将检索和装配表述为一个集合预测问题,并设计了一个新颖的基于Transformer的框架,该框架以变长输出检索零件并回归其姿态。此外,我们利用零件姿态估计与目标分割之间的对偶性,通过联合训练和一个新颖的分割增强优化模块来实现。最后,我们整理了一个包含8万多个形状的大规模数据集,结果表明PART能够泛化到场景布局、图像目标和真实世界扫描。项目页面:此https URL。

英文摘要

3D assembly is fundamental to modern manufacturing and digital content creation. In this paper, we present PART, a unified transformer-based framework for 3D part retrieval and assembly: given a target shape and a part library, PART automatically selects the appropriate parts and predicts their 6-DoF poses to reconstruct the target. While prior work has achieved impressive progress on assembling a pre-defined set of parts, this more practical retrieval-based setting remains largely unexplored. The task faces three key challenges: (i) a combinatorially explosive search space that grows exponentially with library size; (ii) variable-length outputs, as different targets require different numbers of parts; and (iii) continuous 6-DoF pose estimation for part assembly. To address these, we formulate retrieval and assembly as a set prediction problem and design a novel transformer-based framework that retrieves parts and regresses their poses with variable-length output. Additionally, we exploit the duality between part pose estimation and target segmentation through joint training and a novel segmentation-enhanced optimization module. Finally, We curate a large-scale dataset of 80K+ shapes, and the results show that PART generalizes to scene layouts, image targets, and real-world scans. Project Page: https://iambrc.github.io/PART-project-page/.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Tencent(腾讯)
  • Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑