arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CGGS:用于以自我为中心的3D场景生成的一致性增强几何高斯点云渲染

CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation

Zhenyu Sun, Xiaohan Zhang, Qi Liu, Huan Wang

arXiv 2607.03819首次发表:更新:

发表机构

School of Future Technology, South China University of Technology; School of Engineering, Westlake University(华南理工大学未来技术学院; 西湖大学工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对以自我为中心的3D场景生成难题,提出CGGS框架。通过微调多视图潜在扩散模型生成2D内容,利用光流等估计深度产生点云,再经几何细化器优化3D高斯重建,提升视觉质量与几何结构。

AI 中文摘要

由于视图重叠有限和个体视角对场景解释的主导影响,以自我为中心的3D场景生成存在挑战。本文提出CGGS,一个文本到3D的框架,旨在增强3D内容感知并解决以自我为中心的场景生成中的几何失真。首先,通过微调具有一致性增强损失的多视图潜在扩散模型来生成与文本描述一致的高保真2D内容。然后,布局装饰器利用光流和点跟踪对应来估计深度,从而从以自我为中心的2D先验生成密集点云作为粗略布局。在此初始化的基础上,提出几何细化器,通过基于熵的互信息深度损失(MID)结合分层优化方案来增强3D高斯重建,以提高视觉质量和几何结构。综合实验表明,CGGS在生成连贯和准确的文本驱动3D场景方面优于以前的方法。项目页面:此https URL。

英文摘要

Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation. These factors hinder the creation of viewpoint-consistent and semantically aligned visual content, as well as the construction of accurate geometric structures. In this paper, we propose CGGS, a text-to-3D framework aiming to enhance 3D-content-awareness and address geometric distortions in ego-centric scene generation. Firstly, the Ego-centric Generator is proposed by fine-tuning a Multi-View Latent Diffusion Model with consistency-augmented loss to generate consistent, high-fidelity 2D content aligned with textual descriptions. Then, Layout Decorator leverages optical flow and point-track correspondence to estimate depth, therefore producing dense point clouds as coarse layouts from the ego-centric 2D priors. Building on this initialization, Geometric Refiner is proposed to enhance 3D Gaussian reconstruction via an entropy-based Mutual Information Depth Loss (MID) combined with a hierarchical optimization scheme for improving visual quality and geometric structure. Comprehensive experiments demonstrate that CGGS outperforms previous methods in generating coherent and accurate text-driven 3D scenes.

CommentsUpdate format

Journal refIEEE Transactions on Image Processing, vol. 35, pp. 7670-7684, 2026

DOI:10.1109/TIP.2026.3710486

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑