arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.25563cs.CV

CodecSplat: 用于前馈式3D高斯泼溅的超紧凑潜在编码

CodecSplat: Ultra-Compact Latent Coding for Feed-Forward 3D Gaussian Splatting

  • Sun Yat-sen University(中山大学)
  • Peking University(北京大学)
  • Pengcheng Laboratory(鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

Pengpeng Yu, Runqing Jiang, Qi Zhang, Dingquan Li, Jing Wang, Yulan Guo

更新

AI总结:

提出CodecSplat框架,通过将压缩集成到前馈式高斯生成流水线中,利用结构化中间特征表示实现超紧凑场景编码,显著降低存储和传输开销。

AI中文摘要:

尽管前馈式3D高斯泼溅无需逐场景优化即可从稀疏上下文视图重建可渲染的高斯基元,但现有流水线并未提供紧凑的场景表示用于存储或传输。一种自然的解决方案是将现有的3DGS压缩方法应用于生成的高斯基元。然而,这种方法作用于最终的不规则3D表示,且与内部特征到高斯的生成过程解耦,限制了压缩效率。为解决此问题,我们引入了CodecSplat,一种用于前馈式3D高斯泼溅的超紧凑潜在编码框架。CodecSplat首先将中间2D高斯生成特征编码为熵编码的场景比特流。在解码器端,潜在特征被重建并用于预测深度和高斯参数,然后映射到3D高斯基元。注意,通过将压缩集成到前馈式高斯生成流水线中,CodecSplat避免了对不规则3D高斯基元的低效压缩,并允许编解码器利用结构化的中间特征表示。我们在前馈式高斯泼溅骨干网络上实例化了CodecSplat,该网络具有深度引导的多视图特征细化和分层学习特征编解码器。在DL3DV和RealEstate10K数据集上,CodecSplat分别实现了23.56-26.36 dB和24.76-27.05 dB的PSNR,每场景仅需20.00-107.77 KiB和3.37-12.51 KiB。这比压缩前馈式生成的高斯基元大约小一个数量级,同时保持了可控的率失真行为。

英文摘要:

While feed-forward 3D Gaussian splatting reconstructs renderable Gaussian primitives from sparse context views without per-scene optimization, existing pipelines do not provide a compact scene representation for storage or transmission. A natural solution is to apply existing 3DGS compression methods to the generated Gaussian primitives. However, this approach operates on the final irregular 3D representation and is decoupled from the internal feature-to-Gaussian generation process, which limits compression efficiency. To address this, we introduce CodecSplat, an ultra-compact latent coding framework for feed-forward 3D Gaussian splatting. CodecSplat first encodes an intermediate 2D Gaussian-generation feature into an entropy-coded scene bitstream. At the decoder, the latent feature is reconstructed and used to predict depth and Gaussian parameters, which are then mapped to 3D Gaussian primitives. Note that, by integrating compression into the feed-forward Gaussian generation pipeline, CodecSplat avoids inefficient compression over irregular 3D Gaussian primitives and allows the codec to exploit the structured intermediate feature representation. We instantiate CodecSplat on a feed-forward Gaussian splatting backbone with depth-guided multi-view feature refinement and a hierarchical learned feature codec. On DL3DV and RealEstate10K datasets, CodecSplat achieves 23.56-26.36 dB and 24.76-27.05 dB PSNR with only 20.00-107.77 KiB and 3.37-12.51 KiB per scene, respectively. This is roughly one order of magnitude smaller than compressing feed-forward generated Gaussian primitives, while preserving controllable rate-distortion behavior.

↑