发表机构
University of Western Australia; University of Science and Technology Beijing; Monash University(西澳大利亚大学; 北京科技大学; 莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文推出真实世界空间视听理解基准SAVU-Bench及配套诊断集SAVU-Diag,评估发现音频相关空间感知是瓶颈,提出无需训练的基线SAVU-EA可提升空间接地与匹配能力,凸显跨模态空间关系整合的重要性。
AI 中文摘要
空间视听理解要求模型不仅要识别存在的内容,还要识别事件发生的位置以及它们在不同模态间的关联。现有基准通常依赖模拟场景,评估孤立的空间技能,且对失败模式的诊断洞察力有限。我们推出SAVU-Bench,这是一个真实世界基准,可在三个能力级别和七个评估任务中系统评估空间视听理解能力。我们还推出SAVU-Diag,这是一个与场景关联的诊断集,可将推理问题分解为其前提的 grounding(接地)和对齐子任务。在SAVU-Bench上对12个代表性模型的评估显示,视觉空间接地相对成熟,而涉及音频的空间感知仍是主要瓶颈。SAVU-Diag进一步表明,大多数推理错误与这些前提任务的失败同时发生,不过即使前提任务正确解决,推理差距仍然存在。基于这些发现,我们推出SAVU-EA,这是一个无需训练的证据增强基线,可使空间线索更明确。尽管SAVU-EA大幅提升了空间接地和联合匹配能力,但高级空间推理仍然具有挑战性。我们的发现凸显了对稳健的空间音频感知以及跨模态空间关系的更深层次整合的迫切需求。
英文摘要
Spatial audio-visual understanding requires models to recognize not only what is present, but also where events occur and how they relate across modalities. Existing benchmarks often rely on simulated scenes, evaluate isolated spatial skills, and provide limited diagnostic insight into failure modes. We introduce SAVU-Bench, a real-world benchmark that systematically evaluates spatial audio-visual understanding across three capability levels and seven evaluation tasks. We further introduce SAVU-Diag, a scene-linked diagnostic set that decomposes reasoning questions into their prerequisite grounding and alignment sub-tasks. Evaluation of 12 representative models on SAVU-Bench reveals that while visual spatial grounding is relatively mature, spatial perception involving audio remains a primary bottleneck. SAVU-Diag further demonstrates that most reasoning errors co-occur with failures on these prerequisite tasks, though reasoning gaps persist even when prerequisites are correctly resolved. Motivated by these findings, we introduce SAVU-EA, a training-free evidence-augmented baseline that makes spatial cues more explicit. While SAVU-EA substantially improves spatial grounding and joint matching, high-level spatial reasoning remains challenging. Our findings highlight the urgent need for both robust spatial audio perception and deeper integration of cross-modal spatial relations.
CommentsUnder Review