发表机构
University of Washington; University of Texas at Dallas; Hankuk University of Foreign Studies(华盛顿大学; 德克萨斯大学达拉斯分校; 韩国外国语大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出SceneBind全模态场景表示,结合全局语义与对象中心语义空间插槽解决空间结构缺失问题,还提出匹配方案。通过构建新数据集及训练协议进行训练评估,兼容预训练编码器,实现先进检索并能零样本转移到下游任务。
AI 中文摘要
我们提出了SceneBind,一种对现实场景的全模态表示,具有跨视觉、音频和语言的联合语义和3D空间理解。现有全模态编码器在实例级语义(即存在什么)方面表现出色,但往往缺乏明确的空间结构(即其位置)。SceneBind通过将每个场景表示为语义空间实体来解决这一差距,结合全局语义嵌入和以对象为中心的语义空间插槽。我们还提出了SceneBind匹配,一种整合全局场景相似度与对象对齐的语义空间匹配方案,支持跨模态场景检索和对象定位。为训练和评估SceneBind,我们精心构建了一个具有结构化语义和空间注释的新型真实世界双耳视听数据集,并提出了一种跨模态对齐语义和空间信号的训练协议。SceneBind与大规模预训练语义编码器兼容,仅添加少量额外令牌即可进行轻量级空间建模。它实现了最先进的场景和空间检索,同时能够强大地零样本转移到下游任务,如视听定位。
英文摘要
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneBind Matching, a semantic-spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding. To train and evaluate SceneBind, we curate a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, and propose a training protocol for aligning semantic and spatial signals across modalities. SceneBind is compatible with large-scale pretrained semantic encoders, adds lightweight spatial modeling with only a few additional tokens. It achieves state-of-the-art scene and spatial retrieval while enabling strong zero-shot transfer to downstream tasks such as audio-visual localization.
CommentsProject website: https://scenebind.github.io/