发表机构
Nanjing University; Nanjing Agricultural University(南京大学; 南京农业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对野外场景下三维高斯溅射的多视图一致性问题,提出WilLaGS框架,通过β-VAE与三维神经外观场建模外观,结合自监督感知掩码抑制伪影,实现鲁棒三维重建与实时渲染,性能达SOTA。
AI 中文摘要
三维高斯溅射(3DGS)可实现实时高保真渲染,但在无约束的野外场景中仍面临挑战,剧烈的外观变化和瞬态物体破坏了多视图一致性。现有方法受限于独立离散的嵌入,难以捕捉连续环境变化或建模空间变化的局部光照。为解决这些局限,我们提出WilLaGS,这是一种用于无约束场景下鲁棒三维场景重建和生成式外观合成的统一框架。具体而言,我们引入生成式外观模型,其中β-VAE学习结构化连续的全局外观流形;基于隐码,我们构建三维神经外观场,生成动态三平面特征以编码空间变化的局部光照效果。此外,为抑制瞬态伪影,我们提出自监督感知掩码机制,利用教师-学生(EMA)架构推导稳定的场景共识,通过感知差异鲁棒识别不一致区域。在多个数据集上的大量实验表明,WilLaGS在重建质量和新视图外观合成方面达到了最先进性能,同时保持实时渲染效率。
英文摘要
3D Gaussian Splatting (3DGS) delivers real-time and high-fidelity rendering but remains challenged by unconstrained in-the-wild scenes, where drastic appearance variations and transient objects violate multi-view consistency. Existing methods are fundamentally limited by independent and discrete embeddings that struggle to capture continuous environmental changes or model spatially-varying local illumination. To address these limitations, we propose \textbf{WilLaGS}, a unified framework for robust 3D scene reconstruction and generative appearance synthesis under unconstrained settings. Specifically, we introduce a generative appearance model where a $β$-VAE learns a structured and continuous manifold of global appearance. Conditioned on the latent code, we construct a 3D neural appearance field that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects. Furthermore, to suppress transient artifacts, we present a self-supervised perceptual masking mechanism that leverages a Teacher-Student (EMA) architecture to derive a stable scene consensus, robustly identifying inconsistent regions via perceptual discrepancies. Extensive experiments on multiple datasets demonstrate that \textbf{WilLaGS} achieves state-of-the-art performance in reconstruction quality and novel view appearance synthesis, while maintaining real-time rendering efficiency.
CommentsAccepted by ECCV2026