发表机构
The University of Hong Kong; Shanghai Artificial Intelligence Laboratory; Fudan University; Tongji University; University of Science and Technology of China; Shanghai Jiao Tong University(香港大学; 上海人工智能实验室; 复旦大学; 同济大学; 中国科学技术大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MegaParts提出结合结构化序列建模与令牌高效矢量量化形状令牌化器的自回归框架,将部件感知3D生成扩展至300个部件,性能优于基线模型,为大规模部件感知3D生成提供新方案。
AI 中文摘要
部件感知三维物体生成对于可控建模、编辑和关节运动等图形应用至关重要,这类应用中物体被表示为语义部件的连贯组合。然而现有的部件感知生成方法难以很好地扩展到高度复杂的物体:随着部件数量增加,生成详细几何结构所需的令牌长度和内存成本会高得令人望而却步。我们提出MegaParts,一种可扩展的自回归三维生成框架,通过将结构化序列建模与令牌高效的矢量量化形状令牌化器相结合来应对这一挑战。我们的令牌化器在高保真重建的约束下,通过最小化令牌使用量来学习部件级几何结构的离散潜在表示,可基于几何复杂度实现自适应长度的令牌化。在这种紧凑表示的基础上,我们训练了一个大型语言模型(LLM),在统一的结构化序列中生成物体边界框、部件边界框和部件形状令牌。结合高效的长上下文训练策略,我们的令牌高效公式可扩展到包含多达300个部件、序列长度达256k令牌的物体,这在保留组合结构并实现细粒度部件级控制的同时,大幅扩展了部件感知三维生成的规模。我们的方法比基线自回归模型和扩散模型实现了更高的网格质量,表明压缩的离散部件令牌不仅提高了可扩展性,还提升了生成几何结构的可实现保真度。这些结果表明,LLM原生的令牌高效自回归建模是大规模部件感知三维生成的一种极具吸引力的替代扩散方案。项目页面可访问此httpsURL。
英文摘要
Part-aware 3D object generation is essential for graphics applications such as controllable modeling, editing, and articulation, where objects are represented as coherent assemblies of semantic parts. However, existing part-aware generation methods, do not scale well to highly complex objects. As the number of parts increases, generating detailed geometry becomes prohibitively expensive in token length and memory. We introduce MegaParts, a scalable autoregressive 3D generation framework to address this challenge by combining structured sequence modeling with a token-efficient vector-quantized shape tokenizer. Our tokenizer learns discrete latent representations for part-level geometry by minimizing token usage subject to high-fidelity reconstruction, enabling adaptive-length tokenization based on geometric complexity. On top of this compact representation, we train a large language model to generate object bounding boxes, part bounding boxes, and part shape tokens within a unified structured sequence. Combined with efficient long-context training strategy, our token-efficient formulation scales to objects with up to 300 parts and sequence lengths up to 256k tokens. This substantially extends the scale of part-aware 3D generation while preserving compositional structure and enabling fine-grained part-level control. Our method achieves higher mesh quality than baseline autoregressive and diffusion models, showing that compressed discrete part tokens improve not only scalability but also the achievable fidelity of generated geometry. These results suggest that LLM native token-efficient autoregressive modeling is a compelling alternative to diffusion for large-scale part-aware 3D generation. The project page is available at https://expmaster.github.io/megaparts_webpage.
Comments12 pages, 6 pages appendix, 13 figures, technical report