发表机构
Taiyuan University of Technology; Tencent Visvise; Hong Kong University of Science and Technology; Lingnan University; MIT; Macau University of Science and Technology; Texas A&M University(太原理工大学; 腾讯Visvise; 香港科技大学; 岭南大学; 麻省理工学院; 澳门科技大学; 德克萨斯A&M大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PartLLM提出统一多模态模型,将3D部件分割视为意图条件生成问题,通过自回归语义分解支持文本引导、交互及全形状分割,实验证明优于任务特定基线。
AI 中文摘要
部件分割是计算机图形学和3D视觉中的一个基本问题。近期工作已将3D部件分割扩展到固定分类体系之外,但现有方法通常仅针对特定设置,如文本引导的部件分割或基于点的交互。在本工作中,我们认为这些设置可以统一为意图条件生成问题,其中不同的提示指定期望的部件分解方式。为此,我们提出PartLLM,一个统一的多模态模型,将3D部件分割表述为自回归语义分解。在输入形状和用户提示的条件下,PartLLM自回归地生成语义部件假设作为掩码预测的查询,并将其馈送到分解感知解码器,该解码器联合预测连贯的部件掩码。这一统一设计支持文本引导的部件分割、交互式分割以及具有可控粒度的全形状语义分解,且均在单一模型内完成。在这些任务设置上的大量实验表明,PartLLM始终优于任务特定的基线,证明了在意图条件生成框架下统一3D部件分割的有效性。
英文摘要
Part segmentation is a fundamental problem in computer graphics and 3D vision. Recent works have expanded 3D part segmentation beyond fixed taxonomies, but existing approaches typically only address a specific setting, such as text-guided part segmentation or point-based interaction. In this work, we argue that these settings can be unified as an intent-conditioned generative problem, where different prompts specify the desired part decomposition. To this end, we introduce PartLLM, a unified multimodal model that formulates 3D part segmentation as autoregressive semantic decomposition. Conditioned on an input shape and a user prompt, PartLLM autoregressively generates semantic part hypotheses as queries for mask prediction and feeds them to a decomposition-aware decoder that jointly predicts coherent part masks. This unified design supports text-guided part segmentation, interactive segmentation, and full-shape semantic decomposition with controllable granularity within a single model. Extensive experiments across these task settings show that PartLLM consistently outperforms task-specific baselines, demonstrating the effectiveness of unifying 3D part segmentation under an intent-conditioned generative formulation.
CommentsAccepted to SIGGRAPH Asia 2026 (ACM Transactions on Graphics). Project Page: https://czvvd.github.io/PartLLMPage/