AI 中文总结
本文提出首个支持任意分辨率视图的前向网格重建框架AnyGS2Mesh,通过三个关键组件实现高效高质量网格重建,推理速度较优化基线大幅提升,达成近实时效果。
AI 中文摘要
现有的从高斯场景表示进行3D网格重建的方法主要依赖迭代优化,导致推理速度慢,且难以扩展到高分辨率输入。本文提出了AnyGS2Mesh,这是首个支持任意输入图像分辨率、可直接从3D高斯溅射表示重建3D网格的前向框架。该方法采用高斯引导Transformer架构,利用显式3D几何先验实现高效网格生成,包含三个关键组件:(1)高斯引导空间推理Transformer,将高斯基元表示为结构化3D token,共同推理高斯与图像特征;(2)流式分块几何编码器,按顺序处理原生分辨率视图,并聚合可变长度视图集的信息;(3)尺度对齐混合深度细化器,采用PatchFusion风格的编码器-解码器,融合RGB条件预测深度与高斯渲染的度量深度,结合精细局部结构与全局一致的度量尺度。细化后的深度图通过TSDF融合,再经移动立方体算法(Marching Cubes)提取确定性网格。大量实验表明,与基于优化的基线方法相比,AnyGS2Mesh在达到最先进重建质量的同时大幅缩短了推理时间,实现了近实时高质量网格重建,结果证明了高斯表示与前向Transformer架构结合用于可扩展3D几何重建的潜力,代码将在论文接收后公开。
英文摘要
Existing 3D mesh reconstruction methods from Gaussian scene representations predominantly rely on iterative optimization, resulting in slow inference and limited scalability to high-resolution inputs. In this paper, we present AnyGS2Mesh, the first feed-forward framework for directly reconstructing 3D meshes from 3D Gaussian Splatting representations with support for arbitrary input image resolutions. Our approach incorporates a Gaussian-Guided Transformer architecture that exploits explicit 3D geometric priors for efficient mesh generation. We introduce three key components: (1) a Gaussian-Guided Spatial Reasoning Transformer represents Gaussian primitives as structured 3D tokens and jointly reasons over Gaussian and image features; (2) a Streaming and Patchwise Geometry Encoder processes native-resolution views sequentially and aggregates information across variable-length view sets; (3) a Scale-Aligned Hybrid Depth Refiner uses a PatchFusion-style encoder--decoder to fuse RGB-conditioned predicted depth with Gaussian-rendered metric depth, combining fine local structures with globally consistent metric scale. The refined depth maps are integrated through TSDF fusion, followed by Marching Cubes for deterministic mesh extraction. Extensive experiments show that AnyGS2Mesh achieves state-of-the-art reconstruction quality while significantly reducing inference time compared with optimization-based baselines, enabling near-real-time, high-quality mesh reconstruction. Our results demonstrate the potential of combining Gaussian representations and feed-forward Transformer architectures for scalable 3D geometry reconstruction. The code will be made publicly available upon acceptance.