arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

融合感知的直接三维高斯生成与结构化块潜在流

Fusion-Aware Direct 3D Gaussian Generation with Structured Patch Latent Flows

Yizhao Wang, Jingbo Wang, Guantao Zhang

arXiv 2609.27779首次发表:更新:

发表机构

Henan Institute of Science and Technology(河南科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种融合感知的分层高斯块表示与结构感知修正流模型,实现直接、快速且高质量的类别引导三维高斯物体生成,并通过渲染反馈融合提升多视图一致性。

AI 中文摘要

类别引导的三维物体生成对于智能内容创作、虚拟环境和数字资产设计至关重要。尽管三维高斯泼溅(3DGS)提供了显式且渲染高效的表示,但由于高斯原语是无序的、大小可变、局部密集且对渲染高度敏感,直接生成三维高斯物体十分困难。现有的3DGS生成方法通常依赖于多视图合成、重建或提升的2D先验,主要从观察视图中融合信息,而非建模三维高斯物体的内在结构分布。\n本文提出了一种融合感知的分层高斯块表示,用于直接类别引导的3DGS生成,并采用修正流。不规则的高斯集合被分解为规范局部块,并编码为结构化标记。由此产生的分层潜在空间融合了全局类别语义、块级几何与外观、空间对应关系以及渲染敏感线索。在此基础上,我们设计了一种结构感知的修正流模型,具有块位置条件化、全局-局部耦合速度预测和密度感知速度加权,能够在几秒内直接生成类别条件下的3DGS物体。一种渲染反馈融合策略进一步将潜在流学习与解码后的多视图渲染质量对齐。\n实验表明,所提出的方法生成的三维高斯物体具有比基线潜在生成模型更连贯的几何结构、更锐利的局部细节和更好的多视图一致性。消融研究证实了分层信息融合、全局-局部耦合、密度感知监督和渲染反馈学习的贡献,同时保持了实用的整体采样效率。

英文摘要

Class-guided 3D object generation is important for intelligent content creation, virtual environments, and digital asset design. Although 3D Gaussian Splatting (3DGS) offers an explicit and render-efficient representation, directly generating 3D Gaussian objects is difficult because Gaussian primitives are unordered, variable-sized, locally dense, and highly sensitive to rendering. Existing 3DGS generation methods usually depend on multi-view synthesis, reconstruction, or lifted 2D priors, fusing information mainly from observed views rather than modeling the intrinsic structural distribution of 3D Gaussian objects. This paper proposes a fusion-aware hierarchical Gaussian patch representation for direct class-guided 3DGS generation with rectified flow. Irregular Gaussian sets are decomposed into canonical local patches and encoded as structured tokens. The resulting hierarchical latent space fuses global class semantics, patch-level geometry and appearance, spatial correspondence, and rendering-sensitive cues. On this basis, we design a structure-aware rectified flow model with patch-position conditioning, global-local coupled velocity prediction, and density-aware velocity weighting, enabling direct latent generation of class-conditioned 3DGS objects within seconds. A render-feedback fusion strategy further aligns latent flow learning with decoded multi-view rendering quality. Experiments show that the proposed method generates 3D Gaussian objects with more coherent geometry, sharper local details, and better multi-view consistency than baseline latent generative models. Ablation studies confirm the contributions of hierarchical information fusion, global-local coupling, density-aware supervision, and render-feedback learning while preserving practical sampling efficiency overall.

Comments29 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑