arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAC-Splat:用于高保真稀疏视图重建的多属性一致性

MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction

Jinqian Yang, Yichen Wu, Wanhua Li, Haokun Lin, Renzhen Wang, Xiangchu Feng, Xixi Jia

arXiv 2607.10792首次发表:更新:

发表机构

Xidian University; Harvard University; Nanyang Technological University; City University of Hong Kong; Xi’an Jiaotong University(西安电子科技大学; 哈佛大学; 南洋理工大学; 香港城市大学; 西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对稀疏视图重建中现有方法存在几何伪影的问题,提出MAC-Splat训练框架,利用MASt3R和DINOv3获取2D对应关系并定义MAC损失,联合正则化3D属性,实验证明该方法能有效解决不适定的稀疏视图重建问题,性能优于基线。

AI 中文摘要

从稀疏视图重建高保真3D场景一直是可推广神经渲染中的核心问题。现有的可推广3D高斯喷溅(3DGS)方法在稀疏视图设置中常出现几何伪影,因为仅基于2D光度损失的监督无法解决深度和对应模糊性。为此提出MAC-Splat,一个围绕直接3D一致性监督构建的训练框架。它基于MASt3R几何主干和冻结的DINOv3编码器获取语义信息丰富的2D对应关系,以此定义多属性一致性(MAC)损失,联合正则化匹配高斯的3D属性。实验表明MAC-Splat优于强基线,尤其在不同重叠情况下有显著提升,有效解决不适定的稀疏视图重建问题。

英文摘要

Reconstructing high-fidelity 3D scenes from sparse-views remains a central problem in generalizable neural rendering. Existing generalizable 3D Gaussian Splatting (3DGS) methods often exhibit geometric artifacts in sparse-view settings, since supervision based solely on 2D photometric losses cannot resolve depth and correspondence ambiguities. To address this issue, we propose MAC-Splat, a training framework built around direct 3D consistency supervision. MAC-Splat builds on the MASt3R geometric backbone and a frozen DINOv3 encoder to obtain semantically informed 2D correspondences, which serve as geometric anchors for 3D supervision. Using these anchors, we define the Multi-Attribute Consistency (MAC) loss. This objective jointly regularizes the 3D attributes of matched Gaussians, including their position, shape, and appearance, by enforcing agreement in a common world coordinate frame. The formulation is robust to outliers and respects the geometry of covariance matrices, which leads to stable training under sparse-view conditions. Experiments on ScanNet++ show that MAC-Splat outperforms strong baselines, with particularly large gains under different overlap regimes. In particular, it improves average PSNR over Splatt3R by more than 4.5 dB, reduces LPIPS, and maintains performance as the camera pose gap increases. These results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.

CommentsAccepted to the European Conference on Computer Vision (ECCV 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑