AI 中文总结
研究开放词汇3D场景理解中存储瓶颈问题,提出CoSAG方法,通过特定构建方式和存储编码,无需场景训练,实现亚兆字节存储,在多项协议上匹配或超越现有技术,大幅减小场大小。
AI 中文摘要
开放词汇3D场景理解通常通过将诸如CLIP等2D视觉语言特征嵌入3D高斯点云场景,将其转变为可文本查询的语义场来实现。然而,为数百万个高斯点云的每一个都附加高维特征会使单个场景膨胀到千兆字节,这使得存储和部署成为这些领域的真正瓶颈。现有紧凑方法将场构建与场存储纠缠在一起,且从不压缩占成本大部分的每个高斯点云的分配。我们提出CoSAG,它通过闭式透射率加权提升、空间锚定语义锚和多视图去噪在无任何场景训练的情况下构建场,并使用不传输解码器的空间预测熵编码器进行存储。由于锚是空间锚定的,绑定是可预测的,因此高度可压缩。CoSAG在达到亚兆字节存储的同时,在2D渲染、3D选择和密集LSeg协议上匹配或超过现有技术水平,相对于LangSplatV2在更高精度下将场大小缩小3x到76x。
英文摘要
Open-vocabulary 3D scene understanding is commonly achieved by embedding 2D vision-language features such as CLIP into a 3D Gaussian Splatting scene, turning it into a text-queryable semantic field. However, attaching a high-dimensional feature to each of millions of Gaussians inflates a single scene to gigabytes, which makes storage and deployment the real bottleneck of these fields. Existing compact methods each learn and ship a per-scene codec, an autoencoder, a quantized codebook, or a distilled feature field, entangling field construction with field storage and never compressing the per-Gaussian assignment that holds the bulk of the cost. We argue that construction and storage should be decoupled, and that storage is a rate-distortion problem over the per-Gaussian binding to a small anchor table, a structure no prior open-vocabulary method compresses. We present CoSAG, which constructs the field without any per-scene training through a closed-form transmittance-weighted lift, spatially grounded semantic anchors, and multi-view denoising, and stores it with a spatially predictive entropy coder that ships no decoder. Because the anchors are spatially grounded, the binding is predictable and therefore highly compressible. The transmittance-weighted lift and multi-view denoising yield a clean, view-consistent assignment, so the entropy coder spends almost no rate on correcting noise and instead codes only the residual against its spatial prediction. CoSAG reaches sub-megabyte storage while matching or exceeding the state of the art across the 2D-rendered, 3D-selection, and dense-LSeg protocols, reducing field size by 37 to 76x relative to LangSplatV2 at higher accuracy.