Hunyuan3D-Buffalo 1.0:用于可扩展3D生成、理解与编辑的统一多模态模型
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
- Tencent Hunyuan(腾讯混元)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Hunyuan3D-Buffalo 1.0是支持3D理解、文本到3D生成等功能的统一多模态模型,构建87M规模语料库训练,在相关基准上达SOTA或领先性能,验证了统一训练的有效性。
AI中文摘要:
图像生成领域的最新进展已证明,集成理解、生成与编辑功能的统一多模态模型具备潜力。然而,统一3D建模受限于多模态数据稀缺的问题,尤其是缺乏大规模且几何一致性的编辑数据。为解决这一局限,我们提出Hunyuan3D-Buffalo 1.0,这是一个支持3D理解、文本到3D生成、指令引导的3D编辑以及文本 grounding 部件生成的统一框架,所有功能均在单一架构内实现。为实现可扩展训练,我们构建了规模达87M的3D多模态语料库,包含25M个理解样本、50M个文本到3D配对样本,以及使用Nano3D-v2生成的12M个编辑配对样本。在架构上,该框架结合用于语义、结构与空间理解的Hunyuan3D-VLM,以及用于高保真3D合成的Hunyuan3D DiT。VLM为生成提供多模态语义条件,而编辑与部件生成还会对源对象表示的扩散过程施加条件,以保留其整体结构与未编辑区域。大量实验表明,Hunyuan3D-Buffalo 1.0在文本到3D生成与3D编辑基准上达到了SOTA或领先性能,同时展现出强大的理解与部件生成能力。我们的分析进一步显示,生成与理解均能提升编辑效果,证明了统一3D多模态训练的有效性。项目页面:this https URL
英文摘要:
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/