arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2512.00877cs.CV

前馈3D高斯散射压缩与长上下文建模

Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling

  • Hong Kong University of Science and Technology(香港科技大学)
  • Institute of Artificial Intelligence (TeleAI)(人工智能研究所)
  • China Telecom(中国电信)

机构由 AI 辅助整理,请以论文原文为准。

Zhening Liu, Rui Song, Yushi Huang, Yingdong Hu, Xinjie Zhang, Jiawei Shao, Zehong Lin, Jun Zhang

更新

AI总结:

本文提出了一种前馈3DGS压缩方法,通过大规模上下文结构和自回归熵模型实现长距离依赖建模,达到20倍压缩比和先进性能。

AI中文摘要:

3D高斯散射(3DGS)已成为一种革命性的3D表示方法。然而,其庞大的数据量成为广泛应用的主要障碍。虽然前馈3DGS压缩提供了一种替代昂贵的每场景每训练压缩器的实用方案,但现有方法难以建模长距离空间依赖性,由于变换编码网络的有限感受野以及熵模型中不足的上下文容量。在本工作中,我们提出了一种新的前馈3DGS压缩框架,能够有效建模长距离相关性,从而实现高度紧凑且通用的3D表示。我们的方法核心是一个大规模上下文结构,基于Morton序列包含成千上万的高斯分布。我们随后设计了一种细粒度的空间-通道自回归熵模型,以充分利用这一广阔的上下文。此外,我们开发了基于注意力的变换编码模型,通过聚合来自广泛邻近高斯分布的特征来提取信息性的潜在先验。我们的方法在前馈推断中实现了3DGS的20倍压缩比,并在通用编码器中实现了最先进的性能。

英文摘要:

3D Gaussian Splatting (3DGS) has emerged as a revolutionary 3D representation. However, its substantial data size poses a major barrier to widespread adoption. While feed-forward 3DGS compression offers a practical alternative to costly per-scene per-train compressors, existing methods struggle to model long-range spatial dependencies, due to the limited receptive field of transform coding networks and the inadequate context capacity in entropy models. In this work, we propose a novel feed-forward 3DGS compression framework that effectively models long-range correlations to enable highly compact and generalizable 3D representations. Central to our approach is a large-scale context structure that comprises thousands of Gaussians based on Morton serialization. We then design a fine-grained space-channel auto-regressive entropy model to fully leverage this expansive context. Furthermore, we develop an attention-based transform coding model to extract informative latent priors by aggregating features from a wide range of neighboring Gaussians. Our method yields a $20\times$ compression ratio for 3DGS in a feed-forward inference and achieves state-of-the-art performance among generalizable codecs.

↑