发表机构
VCIP, Nankai University; Tencent Hunyuan; MMLab, CUHK; Fudan University; Shanghai Innovation Institute; HKU(南开大学VCIP; 腾讯混元; 香港中文大学MMLab; 复旦大学; 上海创新研究院; 香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对单阶段三维生成模型缺乏显式位置引导的问题,提出位置强制框架,通过去噪中恢复并反馈逐步精细的标记位置,实现从粗到细的自条件生成,显著提升生成质量并超越多种多阶段方法。
AI 中文摘要
近期单阶段三维生成模型普遍采用VecSet表示,将三维形状编码为潜在标记的无序集合。然而,与提供显式位置引导的两阶段方法相比,这些模型必须在去噪过程中隐式推断标记位置,从而限制了其生成质量。我们观察到,尽管缺乏显式位置条件,VecSet标记仍保留可恢复的空间对应关系。基于这一观察,我们提出位置强制(Position Forcing),一种基于位置的自条件框架。在去噪过程中,位置强制从当前干净潜在估计中恢复标记位置,根据去噪阶段以逐步精细的分辨率对其进行量化,并将所得位置编码反馈回扩散Transformer。这种逐步精细的位置反馈提供了与每个去噪阶段相适应的粒度空间引导,沿从粗到细的轨迹引导形状生成,在无需单独位置生成阶段的情况下大幅提升生成质量。实验表明,位置强制在单阶段三维生成方法中取得了强劲性能,并超越了多种有竞争力的多阶段方法。
英文摘要
Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation quality. We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences. Building on this observation, we propose Position Forcing, a position-based self-conditioning framework. During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer. This progressively refined positional feedback provides spatial guidance at a granularity appropriate to each denoising stage, guiding shape generation along a coarse-to-fine trajectory and substantially improving generation quality without a separate position generation stage. Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.