基于自引导稀疏体细化的精细细节单目几何估计
MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement
- Tsinghua University(清华大学)
- USTC(中国科学技术大学)
- Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对单目几何估计在局部3D结构精细细节上的失真问题,提出基于自引导稀疏体细化的方法,将建模从2D提升到3D空间,通过稀疏卷积避免特征混合,实验证明该方法在恢复精细3D几何方面显著优于现有方法。
AI中文摘要:
单目几何估计在不同场景中取得了显著性能,但当前最先进模型在局部3D结构尤其是精细细节上仍有明显失真。我们将此局限归因于架构不匹配,多数模型在2D参数化内解码3D几何,导致特征混合。本文提出自引导稀疏体细化(SSR)的精细细节单目几何估计,将单目几何建模从2D图像空间提升到3D空间。模型将基础模型的粗点图提升到稀疏体素壳上并通过SSR细化,SSR采用基于3D空间局部性聚合特征的稀疏卷积。实验表明该方法在恢复精细3D几何上显著优于现有方法。
英文摘要:
Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structure, especially in fine details, like thin structures and small objects. We attribute this limitation to an architectural mismatch: most current models decode 3D geometry within a 2D parameterization, where feature interactions are governed by image-plane proximity rather than true 3D spatial relationships. This inadvertently mixes features from geometrically distant surfaces, resulting in over-smoothed geometry particularly around thin or elongated structure. In this paper, we propose MoGe-3, a fine-detail monocular geometry estimation model with Self-Guided Sparse 3D Refinement (SSR) that lifts monocular geometry modeling from 2D image space to 3D space for high-fidelity metric-scale point maps. MoGe-3 lifts the coarse point map from a foundation base model onto a sparse voxel shell and refines it via SSR. The SSR employs sparse convolutions that aggregate features based on 3D spatial locality, avoiding feature mixing across depth discontinuities. Extensive experiments on diverse datasets demonstrate that MoGe-3 significantly outperforms existing approaches in recovering fine detailed 3D geometry across both quantitative metrics and qualitative visualizations. Project page: https://qft-333.github.io/moge3page/