arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SparseTalk - 稀疏化3D高斯语言场以实现高效3D视觉问答

SparseTalk - Sparsifying 3D Gaussian Language Fields for Efficient 3D Visual Question Answering

Davit Soselia, Joseph JaJa, Amitabh Varshney

arXiv 2609.15137首次发表:更新:

发表机构

University of Maryland(马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SparseTalk通过对象级稀疏化将3D高斯语言场压缩至不足1%,在ScanQA和MV-ScanQA上保持强VQA性能,推理输入减少至0.80%,解码内存降低125倍。

AI 中文摘要

3D高斯语言场为3D视觉问答(VQA)提供了显式、空间锚定的表示,但其密集的语义特征每个场景可能需要数万个嵌入,导致大量的存储、内存和推理成本。我们研究了这种表示中实际有多少对于下游推理是必要的。从完整的嵌入表示出发,我们系统地稀疏化其语义嵌入,包括先前未充分探索的低于单个图像等效块直至8个视觉标记的区间。我们比较了随机、几何、语义和联合空间-语义选择策略,并引入了一种基于对象的稀疏化方法,该方法在检测到的对象实例之间分配标记预算,同时保留背景上下文。在ScanQA和MV-ScanQA上的实验揭示了密集高斯语言场中存在大量冗余。仅需几百个语义嵌入(相当于原始表示的不到1%)即可保持强大的VQA性能。基于对象的选择相对于其他方法表现良好,在降至256个标记时仅观察到适度的变化。在此预算下,SparseTalk保留了SplatTalk的32,076个标记推理输入的0.80%和平均77,207个高斯密集场的0.332%,同时提高了推理吞吐量并将解码特征内存减少了125倍。

英文摘要

3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic features can require tens of thousands of embeddings per scene, resulting in substantial storage, memory, and inference costs. We investigate how much of this representation is actually necessary for downstream reasoning. Starting from a full embedding representation, we systematically sparsify its semantic embeddings, including the previously underexplored regime below a single image-equivalent block down to 8 visual tokens. We compare random, geometric, semantic, and joint spatial-semantic selection strategies and introduce an object-based sparsification method that distributes the token budget across detected object instances while retaining background context. Experiments on ScanQA and MV-ScanQA reveal substantial redundancy in dense Gaussian language fields. Strong VQA performance is retained with only a few hundred semantic embeddings, corresponding to less than 1% of the original representation. Object-based selection performs well relative to others, with only modest observed changes down to 256 tokens. At this budget, SparseTalk retains 0.80% of SplatTalk's 32,076-token inference input and 0.332% of the mean 77,207-Gaussian dense field, increasing inference throughput while reducing decoded-feature memory 125-fold.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑