arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24981cs.CV

GAE:学习几何原生的潜在空间用于3D一致的世界生成

GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出几何原生自编码器(GAE),将几何基础模型特征重参数化为紧凑潜在空间,联合解码外观、深度、相机和点图,在保持生成器不变下提升3D一致性和视觉质量,FVD显著下降。

中文摘要 AI 辅助

我们提出了一个紧凑的几何原生潜在空间,作为感知和生成的共享基础。视觉生成器可以生成逼真的帧,而无需保持一致的3D场景。我们认为这不仅是一个建模问题,也是一个表示问题:生成器通常演化以外观为中心的潜在表示,而感知模型在语义丰富的空间中恢复几何结构,该空间编码跨视图结构。我们没有将几何作为另一个输出添加,而是将几何基础模型的特征重新参数化为一个紧凑的潜在空间用于生成。我们通过几何原生自编码器(GAE)实现了这一转变,其潜在表示可联合解码为外观、深度、相机和点图。利用这种状态,一个标准的条件流支持多种生成任务。在保持生成器和训练协议固定的受控比较中,用GAE替换潜在表示提高了视觉质量和独立测量的3D一致性:在RealEstate10K和DL3DV上,FVD分别下降了12.7%和23.1%,在RealEstate10K上相机轨迹误差减半。这些结果共同表明,潜在空间对于几何一致的生成至关重要,并且可以作为感知和生成之间的共享接口。

英文摘要

We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by $12.7\%$ and $23.1\%$ on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.

发表机构

  • The Hong Kong University of Science and Technology(香港科技大学)
  • ARC Lab, Tencent IEG(腾讯互动娱乐事业群 ARC 实验室)
  • The University of Hong Kong(香港大学)
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑