arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PartLLM:面向3D部件分割的统一多模态基础模型

PartLLM: A Unified Multimodal Foundation for 3D Part Segmentation

Zhe Zhu, Yiheng Zhang, Peng Li, Zixing Zhao, Honghua Chen, Yaqing Zhang, Le Wan, Zhiyang Dou, Cheng Lin, Yuan Liu, Mingqiang Wei, Wenping Wang

arXiv 2609.25832首次发表:更新:

发表机构

Taiyuan University of Technology; Tencent Visvise; Hong Kong University of Science and Technology; Lingnan University; MIT; Macau University of Science and Technology; Texas A&M University(太原理工大学; 腾讯Visvise; 香港科技大学; 岭南大学; 麻省理工学院; 澳门科技大学; 德克萨斯A&M大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PartLLM提出统一多模态模型,将3D部件分割视为意图条件生成问题,通过自回归语义分解支持文本引导、交互及全形状分割,实验证明优于任务特定基线。

AI 中文摘要

部件分割是计算机图形学和3D视觉中的一个基本问题。近期工作已将3D部件分割扩展到固定分类体系之外,但现有方法通常仅针对特定设置,如文本引导的部件分割或基于点的交互。在本工作中,我们认为这些设置可以统一为意图条件生成问题,其中不同的提示指定期望的部件分解方式。为此,我们提出PartLLM,一个统一的多模态模型,将3D部件分割表述为自回归语义分解。在输入形状和用户提示的条件下,PartLLM自回归地生成语义部件假设作为掩码预测的查询,并将其馈送到分解感知解码器,该解码器联合预测连贯的部件掩码。这一统一设计支持文本引导的部件分割、交互式分割以及具有可控粒度的全形状语义分解,且均在单一模型内完成。在这些任务设置上的大量实验表明,PartLLM始终优于任务特定的基线,证明了在意图条件生成框架下统一3D部件分割的有效性。

英文摘要

Part segmentation is a fundamental problem in computer graphics and 3D vision. Recent works have expanded 3D part segmentation beyond fixed taxonomies, but existing approaches typically only address a specific setting, such as text-guided part segmentation or point-based interaction. In this work, we argue that these settings can be unified as an intent-conditioned generative problem, where different prompts specify the desired part decomposition. To this end, we introduce PartLLM, a unified multimodal model that formulates 3D part segmentation as autoregressive semantic decomposition. Conditioned on an input shape and a user prompt, PartLLM autoregressively generates semantic part hypotheses as queries for mask prediction and feeds them to a decomposition-aware decoder that jointly predicts coherent part masks. This unified design supports text-guided part segmentation, interactive segmentation, and full-shape semantic decomposition with controllable granularity within a single model. Extensive experiments across these task settings show that PartLLM consistently outperforms task-specific baselines, demonstrating the effectiveness of unifying 3D part segmentation under an intent-conditioned generative formulation.

CommentsAccepted to SIGGRAPH Asia 2026 (ACM Transactions on Graphics). Project Page: https://czvvd.github.io/PartLLMPage/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑