发表机构
Huazhong University of Science and Technology; Megvii; Zhejiang University(华中科技大学; 旷视科技; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ROAD框架通过迁移判别式3D基础模型的先验,采用互目标对齐策略,仅用1.5%训练数据就实现了高保真3D生成,大幅降低了计算开销。
AI 中文摘要
高保真3D生成主要依赖于扩大模型容量和数据规模,这会带来高昂的计算成本。这种范式通常需要从头学习几何结构,却忽略了判别式3D基础模型中已包含的丰富语义和结构先验。我们认为,利用这些判别式模型对3D世界的深刻理解可显著降低生成成本。为此,我们提出了ROAD框架,该框架通过将丰富的判别式先验迁移到扩散Transformer中,降低了3D生成的训练成本。为解决生成式与判别式隐空间之间固有的语义-结构异质性,我们引入了互目标对齐策略。该方法协同整体语义压缩以确保全局语义一致性,以及结构最优对齐——其被公式化为二分匹配问题,以严格对齐不同隐空间之间的微观几何细节。3D基础模型仅用于对齐的训练时监督,推理时不使用,因此不会产生额外的推理成本。与工业基线Step1X-3D相比,所提出的ROAD仅用1.5%的训练数据就实现了极具竞争力的生成性能,且显著降低了训练成本,有效减少了高保真3D生成的计算开销。代码可在该https URL获取。
英文摘要
High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typically requires learning geometry from scratch and overlooks the rich semantic and structural priors already encapsulated in discriminative 3D foundation models. We contend that leveraging the profound understanding of the 3D world possessed by these discriminative models can significantly reduce generative cost. To this end, we propose ROAD, a framework that reduces the training cost of 3D generation by transferring these rich discriminative priors into diffusion transformers. To address the inherent semantic-structural heterogeneity between generative and discriminative latents, we introduce a reciprocal-objective alignment strategy. This method synergizes Holistic Semantic Condensing to enforce global semantic coherence and Structural Optimal Alignment, which is formulated as a bipartite matching problem to rigorously align microscopic geometric details between disparate latent spaces. The 3D foundation model is only used for training-time supervision of alignment and is not used at inference, incurring no additional inference cost. Compared with the industrial baseline Step1X-3D, the proposed ROAD achieves highly competitive generation performance with only 1.5% of the training data and significantly reduces training costs, effectively reducing the computational overhead of high-fidelity 3D generation. Code is available at https://github.com/H-EmbodVis/ROAD.