LangStreet:用于锚点解码街道高斯分布的持久语言场
LangStreet: Persistent Language Fields for Anchor-Decoded Street Gaussians
浏览论文内容
中文总结 AI 辅助
针对锚点解码高斯场景中语言场跨视角失效问题,提出持久语言场方法,通过语义所有权和证据累积,在保持精度的同时大幅降低存储成本。
中文摘要 AI 辅助
语言高斯场隐含地假设携带语义的原语在不同视角下保持可识别性。这一假设在可扩展的锚点解码表示中不成立,因为持久锚点会生成依赖于视角的子高斯分布,其几何和外观随相机变化而变化。我们提出了Ours,一种用于此类结构化高斯场景的持久语言场。我们的核心思想是语义所有权:瞬态子节点路由观测,而持久的解码器槽及其父锚点拥有语言场。我们利用alpha合成责任在槽处累积加性方向证据;这些统计量精确边缘化到锚点。然后,我们用锚点对齐的证据补全支持较弱的槽,同时保留锚点方向,并通过锚点相对语义坐标中的低秩残差来表示槽细节。我们的主要模型Ours (base)存储锚点特征以及紧凑的槽残差。Ours (light)仅保留锚点特征,而Ours (max)显式存储全维补全的槽特征。在没有场景特定语义优化的情况下,Ours (base)在KITTI、Virtual KITTI和Waymo上几乎与Ours (max)匹配。在KITTI上,它以2.72 GiB的有效特征占用实现了34.19的2D mIoU,而Ours (max)为34.20 mIoU和12.90 GiB。相同的精度-存储趋势在Virtual KITTI和Waymo上也成立。这些结果表明,基于视角条件样条的语言场需要持久的语义所有权、守恒的证据以及平衡稳定性、细节和表示成本的层次结构。我们的代码、检查点和基准套件将公开提供。
英文摘要
Language Gaussian fields implicitly assume that the primitive carrying semantics remains identifiable across views. This assumption breaks in scalable anchor-decoded representations, where persistent anchors generate view-conditioned child Gaussians whose geometry and appearance vary with the camera. We introduce Ours, a persistent language field for such structured Gaussian scenes. Our key idea is semantic ownership: transient children route observations, while persistent decoder slots and their parent anchors own the language field. We use alpha-compositing responsibilities to accumulate additive directional evidence at slots; these statistics marginalize exactly to anchors. We then complete weakly supported slots with anchor-aligned evidence while preserving the anchor direction, and represent slot detail through low-rank residuals in anchor-relative semantic coordinates. Our primary model, Ours (base), stores anchor features together with compact slot residuals. Ours (light) retains only anchor features, whereas Ours (max) stores the full-dimensional completed slot features explicitly. Without scene-specific semantic optimization, Ours (base) nearly matches Ours (max) across KITTI, Virtual KITTI, and Waymo. On KITTI, it achieves 34.19 2D mIoU with a 2.72 GiB effective feature footprint, compared with 34.20 mIoU and 12.90 GiB for Ours (max). The same accuracy-storage trend holds on Virtual KITTI and Waymo. These results show that language fields on view-conditioned splats require persistent semantic ownership, conserved evidence, and a hierarchy that balances stability, detail, and representation cost. Our code, checkpoints, and benchmark suite will be publicly available.
发表机构
- INSAIT, Sofia University “St. Kliment Ohridski”(索非亚大学INSAIT研究所)
- vivo Mobile Communication Co., Ltd., Shenzhen, China(维沃移动通信有限公司)
机构由 AI 辅助整理,请以论文原文为准。