Block3D:基于分块扩散的高效文本到3D生成
Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion
- ZipLab, Zhejiang University(浙江大学ZipLab)
- University of California, Berkeley(加州大学伯克利分校)
- Wuhan University(武汉大学)
- Monash University(莫纳什大学)
- University of Adelaide(阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对文本到3D生成成本高的问题,提出Block3D分块扩散框架,通过分块生成、联合去噪及置信度引导的块内修正,在保持几何保真度的同时实现5.15倍的速度提升。
中文摘要 AI 辅助
尽管文本到3D生成技术已快速发展,但在低推理成本下实现高几何保真度仍具挑战性。现有文本到3D方法要么自回归解码离散形状token,要么用扩散或流模型迭代优化全局3D表示。然而自回归解码是顺序的,无法修正错误;而扩散和流匹配模型会重复处理完整表示,导致高质量生成的成本越来越高。本文提出Block3D,一种分块扩散框架,它将离散形状token序列划分为连续块,自回归生成各块,并联合去噪当前块内的所有token。为缓解错误累积,我们引入置信度引导的块内修正,在每个块最终确定前修正低置信度token。在TRELLIS-500K的保留集上,Block3D将平均端到端生成时间从25.71秒降至4.99秒,比微调后的自回归基线快5.15倍,且未牺牲几何保真度。
英文摘要
While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains challenging. Existing text-to-3D methods either decode discrete shape tokens autoregressively or iteratively refine global 3D representations with diffusion or flow models. However, autoregressive decoding is sequential and cannot revise errors, whereas diffusion and flow-matching models repeatedly process the full representation, making high-quality generation increasingly expensive. In this paper, we propose Block3D, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block. To alleviate error accumulation, we introduce confidence-guided intra-block correction, which revises low-confidence tokens before each block is finalized. On a held-out set from TRELLIS-500K, Block3D reduces mean end-to-end generation time from 25.71 seconds to 4.99 seconds, achieving a $5.15\times$ speedup over the fine-tuned autoregressive baseline without sacrificing geometric fidelity.