arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01056cs.CVeess.IV

HierGF:基于几何感知消息传递的层次高斯场用于稀疏视图三维重建

HierGF: Hierarchical Gaussian Fields via Geometry-perception Message Passing for Sparse-view 3D Reconstruction

Bi'an Du, Zhimin Zhang, Daizong Liu, Baoquan Chen, Wei Hu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出HierGF,通过层次几何感知消息传递将有限观测转化为伪监督,增强稀疏视图三维重建的多视图一致性与结构完整性。

中文摘要 AI 辅助

稀疏视图三维重建是多媒体应用中的一个重要且常见的场景,例如增强现实/虚拟现实(AR/VR)内容创作、文化遗产数字化以及某些机器人应用,在这些场景中可能只有有限数量的随机捕获视图可用。然而,稀疏视图仅包含有限的三维信息,带来了两大挑战:1)可用于匹配的图像太少,难以建立多视图一致性;2)视图覆盖不足导致欠采样区域缺乏信息,造成物体结构缺失。现有方法大多仍依赖有限的重投影误差和正则化项,这容易过度拟合单一视图并导致跨视图外观不一致。在几何欠采样区域,它们通常依赖启发式密度控制,缺乏可靠指导,常常导致模糊和结构错误。为解决这些问题,本文提出层次高斯场(HierGF),从层次几何感知的角度重新审视稀疏视图重建,并将有限的观测转化为可靠的自我生成监督,超越固定先验和启发式密度控制。具体而言,我们通过两阶段几何感知骨干网络将粗略的三维几何信息和额外的二维生成先验转化为结构化的伪监督,从而在极少的输入视图下增强多视图一致性。此外,我们引入一个可学习的置信度网络,将梯度引导至跨视图一致的内容,并引入一个几何一致的稠密化模块,以改善多视图对齐和欠采样区域的重建。

英文摘要

Sparse view 3D reconstruction is an important and common scenario in multimedia applications, such as augmented reality/virtual reality (AR/VR) content creation, cultural heritage digitization, and certain robotic applications, where only a limited number of randomly captured views may be available. However, sparse views contain only limited 3D information, posing two major challenges:1) too few images are available for matching, making it difficult to build multi-view consistency; 2) insufficient view coverage leads to a lack of information in under-sampled regions, resulting in missing parts of object structure. Existing methods mostly still rely on limited reprojection errors and regularization terms, which are prone to overfitting to a single view and inconsistent appearances across views. In geometrically under-sampled regions, they often rely on heuristic density control, lacking reliable guidance and often resulting in blurring and structural holes.To address these issues, this paper proposes Hierarchical Gaussian Fields (HierGF), which revisits sparse-view reconstruction from a hierarchical geometry-perception perspective and converts limited observations into reliable self-generated supervision beyond fixed priors and heuristic density control. In particular, we transform coarse 3D geometric information and additional 2D generative priors into structured pseudo-supervision through a two-stage geometry-perception backbone network, thereby enhancing multi-view consistency with very few input views. In addition, we introduce a learnable confidence network to guide gradients toward cross-view consistent content, and a geometrically consistent densification module to improve the reconstruction of multi-view alignment and under-sampled regions.

发表机构

  • Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
  • Institute for Math & AI, Wuhan University(武汉大学数学与人工智能研究院)
  • School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑