arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27778cs.CV

可见性引导的结构化测度流用于类别条件三维高斯生成

Visibility-Guided Structured Measure Flow for Class-Conditioned 3D Gaussian Generation

  • Henan Institute of Science and Technology(河南科技学院)

机构由 AI 辅助整理,请以论文原文为准。

Yizhao Wang

AI总结:

VISTA-GS提出可见性引导的结构化测度流,将3DGS对象建模为加权高斯测度,通过测度VAE和渲染器一致的流实现类别条件生成,在VISTA-Obj30上性能提升60-72%。

AI中文摘要:

三维高斯泼溅(3DGS)已使实时、高保真三维渲染成为现实,但将这种显式表示转化为原生生成空间仍是一个开放挑战。直接生成3DGS对象是困难的,因为高斯基元是无序的、可变大小的、局部密集的,并且对渲染行为高度敏感。我们提出了VISTA-GS,一个用于类别条件三维高斯生成的可见性引导结构化测度流框架。我们不将3DGS对象视为平坦的基元序列或通用潜在标记网格,而是将其公式化为由不透明度、各向异性协方差和多视角可见性加权的结构化高斯测度。基于此公式,我们引入了一个可见性感知的测度VAE,学习3DGS对象的排列不变、可变大小兼容且渲染感知的潜在表示。我们进一步开发了一个渲染器一致的测度流,将类别条件先验传输到学习到的3DGS测度分布,同时将解码对象与其多视角渲染分布对齐。为了保留对象布局和局部细节,VISTA-GS结合了结构保持的补丁传输,在流预测期间耦合全局类别语义、局部高斯测度补丁和空间锚点。在VISTA-Obj30上,VISTA-GS在几何、外观、视角一致性误差和生成速度方面比最强基线提高了约60-72%。该设计能够高效生成连贯、详细且视角一致的三维高斯对象,无需依赖逐实例优化、多视角图像合成或基于重建的提升流程。项目代码和模型检查点将发布。

英文摘要:

3D Gaussian Splatting (3DGS) has made real-time, high-fidelity 3D rendering practical, yet turning this explicit representation into a native generative space remains an open challenge. Directly generating 3DGS objects is difficult because Gaussian primitives are unordered, variable-sized, locally dense, and highly sensitive to rendering behavior. We present VISTA-GS, a visibility-guided structured measure flow framework for class-conditioned 3D Gaussian generation. Instead of treating a 3DGS object as a flat primitive sequence or a generic latent token grid, we formulate it as a structured Gaussian measure weighted by opacity, anisotropic covariance, and multi-view visibility. Based on this formulation, we introduce a visibility-aware measure VAE that learns permutation-invariant, variable-size-compatible, and rendering-aware latent representations of 3DGS objects. We further develop a renderer-consistent measure flow that transports class-conditioned priors toward the learned 3DGS measure distribution while aligning the decoded objects with their multi-view rendering distributions. To preserve object layout and local details, VISTA-GS incorporates structure-preserving patch transport that couples global class semantics, local Gaussian measure patches, and spatial anchors during flow prediction. On VISTA-Obj30, VISTA-GS improves over the strongest baseline by roughly 60--72\% across geometry, appearance, view-consistency error, and generation speed. This design enables efficient generation of coherent, detailed, and view-consistent 3D Gaussian objects without relying on per-instance optimization, multi-view image synthesis, or reconstruction-based lifting pipelines. Project code and model checkpoints will be released.

补充信息

↑