发表机构
Cooperative Medianet Innovation Center, Shanghai Jiao Tong University(上海交通大学协同媒体网络创新中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对前馈3D高斯点云压缩难题,提出GenSplatCodec统一编解码器,采用双流编码与几何引导生成解码,经三阶段优化策略,在多数据集实验中展现出优于现有方法的率失真性能。
AI 中文摘要
前馈3D高斯点云(3DGS)可实现可扩展的场景重建,无需逐场景优化,但会产生密集的高斯点云,存储和传输成本高昂。现有前馈高斯压缩方法将解码视为确定性表示恢复,在低比特率下丢弃高频纹理和视图相关外观时变得不足。虽然生成模型提供了有前景的替代方案,但将它们用作独立后处理会使生成与传输的场景结构解耦,从而损害跨视图一致性。为解决这些限制,我们提出GenSplatCodec,一种统一的前馈高斯编解码器,将低比特率高斯压缩重新表述为几何引导的生成解码。我们在双流框架内提出了一种细节感知的前馈高斯编码方案,紧凑的高斯结构流由轻量级参考外观流补充。我们还引入了几何引导的一步生成解码方法,通过分层几何控制联合利用解码的结构和外观线索来重建高保真和视图一致的新视图。最后,我们开发了一种三阶段优化策略,稳定统一编解码器的学习,并使生成解码器适应编解码器衍生的结构和外观线索。跨多个数据集的大量实验表明,GenSplatCodec始终优于现有方法。
英文摘要
Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are costly to store and transmit. Existing feed-forward Gaussian compression methods formulate decoding as deterministic representation recovery, which becomes inadequate at low bitrates when high-frequency textures and view-dependent appearance are discarded. Although generative models offer a promising alternative, using them as standalone post-processing decouples generation from the transmitted scene structure, thereby compromising cross-view consistency. To address these limitations, we propose GenSplatCodec, a unified feed-forward Gaussian codec that reformulates low-bitrate Gaussian compression as geometry-guided generative decoding. We present a detail-aware feed-forward Gaussian coding scheme within a dual-stream formulation, where the resulting compact Gaussian structural stream is complemented by a lightweight reference appearance stream. We further introduce a geometry-guided one-step generative decoding approach that jointly exploits decoded structural and appearance cues through hierarchical geometry control to reconstruct high-fidelity and view-consistent novel views. Finally, we develop a three-stage optimization strategy that stabilizes the learning of the unified codec and adapts the generative decoder to codec-derived structural and appearance cues. Extensive experiments across multiple datasets demonstrate that GenSplatCodec consistently achieves superior rate-distortion (RD) performance over existing methods.