OmniEcho:具身智能体的空间音频理解
OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents
浏览论文内容
中文总结 AI 辅助
本文提出OmniEchoBench基准和OmniEcho模型,用于具身智能体的空间音频理解,通过可控渲染和FOA编码器,在感知和导航任务上达到先进性能,证明空间音频的价值。
中文摘要 AI 辅助
人类可以毫不费力地定位声源的方向,并将其与视觉线索相结合进行推理,然而这对具身智能体而言仍然具有挑战性。特别是,如何在具身场景中有效评估和建模空间音频理解仍不清楚。为解决这一空白,我们引入了\ extbf{OmniEchoBench},一个用于空间音视频感知和音频-视觉-语言导航的统一基准。OmniEchoBench包含六个任务,涵盖197个真实世界空间音视频场景、2,972个问答对,以及从30个真实世界环境中收集的包含一阶环境立体声(FOA)音频的900个导航样本。为了实现可扩展的训练监督,我们开发了一个可控的空间音频渲染流程。该流程保持了声源、视觉观察和智能体轨迹之间的几何一致性。在此基础上,我们提出了\ extbf{OmniEcho},一个具有空间感知能力的全模态模型。它引入了一个FOA空间编码器以及一个预训练的语义音频通路。大量实验表明,OmniEcho在空间音视频感知方面达到了最先进的性能。对于我们的声音引导导航,OmniEcho达到了接近传统视觉-语言导航的性能水平。这些结果表明,空间音频可以作为具身场景推理和导航的有价值信号,同时也突出了细粒度空间定位和距离估计作为重要的开放挑战。
英文摘要
Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce \textbf{OmniEchoBench}, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics (FOA) audio collected from 30 real-world environments. To enable scalable training supervision, we develop a controllable rendering pipeline for spatial audio. It preserves geometric consistency among sound sources, visual observations, and agent trajectories. Building on this, we propose \textbf{OmniEcho}, a spatially aware omni-modal model. It introduces an FOA spatial encoder alongside a pretrained semantic audio pathway. Extensive experiments show that OmniEcho achieves state-of-the-art performance on spatial audio-visual perception. For our sound-guided navigation, OmniEcho reaches a performance level close to that of traditional vision-language navigation. These results demonstrate that spatial audio can serve as a valuable signal for embodied scene reasoning and navigation, while also highlighting fine-grained spatial localization and distance estimation as important open challenges. Our code and data will be available in https://github.com/PKU-VaLuE-Lab/OmniEcho/tree/main
发表机构
- Peking University(北京大学)
- Alibaba Group(阿里巴巴集团)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。