发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3DGS渲染中排序开销和随机噪声问题,提出混合点绘与高斯感知时空重建网络,实现免排序高效渲染,兼顾速度与质量,并支持移动端交互式应用。
AI 中文摘要
传统的3D高斯泼溅(3DGS)需要对高斯图元进行深度排序和有序alpha混合,以正确渲染重叠的高斯图元。随机透明度通过用离散的随机可见性样本替代分数alpha贡献,实现了免排序渲染,但在低样本数下会产生大量的空间和时间噪声。我们将这种从连续高斯泼溅到离散可见性样本的转换称为“高斯点绘”。基于此,我们提出了一种高效的顺序无关渲染与重建框架,可直接处理未经修改的3DGS资产。我们的方法自适应地整合了基于图元和基于片段的点绘,利用它们在不同渲染场景下的互补优势,显著提升了渲染吞吐量。为了从稀疏的随机样本中恢复高质量图像,我们进一步引入了一个轻量级的高斯感知时空重建网络。通过利用每个点绘保留的高斯属性,该网络在空间和时间上聚合结构化的随机线索,有效抑制点绘噪声。实验表明,我们的混合高斯点绘方法,结合在多样场景上训练的时空重建网络,能够泛化到未见过的场景,并在无需重新训练或预处理的情况下,在移动设备上实现交互式、时间稳定且视觉合理的渲染。通过场景特定训练以及适当缩放的采样和网络容量,我们的方法在视觉质量上进一步超越了基线,为质量优先的应用提供了高保真配置。
英文摘要
Conventional 3D Gaussian Splatting (3DGS) requires depth sorting and ordered alpha blending to correctly render overlapping Gaussian primitives. Stochastic transparency enables sorting-free rendering by replacing fractional alpha contributions with discrete stochastic visibility samples, but produces substantial spatial and temporal noise at low sample counts. We refer to this conversion from continuous Gaussian splats to discrete visibility samples as \textit{Gaussian Stippling}. Based on this, we present an efficient order-independent rendering and reconstruction framework that operates directly on unmodified 3DGS assets. Our method adaptively integrates primitive-based and fragment-based stippling, leveraging their complementary strengths across different rendering regimes to significantly improve rendering throughput. To recover high-quality images from sparse stochastic samples, we further introduce a lightweight Gaussian-aware spatiotemporal reconstruction network. By exploiting the Gaussian attributes retained by each stipple, the network aggregates structured stochastic clues across both space and time, effectively suppressing stippling noise. Experiments show that our hybrid Gaussian stippling method, coupled with a spatiotemporal reconstruction network trained on diverse scenes, generalizes to unseen scenes and enables interactive, temporally stable, and visually plausible rendering on mobile devices without retraining or preprocessing. With scene-specific training and appropriately scaled sampling and network capacity, our method further outperforms the baselines in visual quality, offering a high-fidelity configuration for quality-prioritized applications.
CommentsPreprint