RMS-AQA:面向真实世界家庭环境的双阶段空间音频问答基准
RMS-AQA: A Two-Stage Spatial Audio Question Answering Benchmark for Real-World Domestic Environments
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对家庭具身助手的空间音频理解需求,提出双阶段SAQA基准RMS-AQA,结合真实FOA录音与RIR合成数据,并引入轻量空间插件,实验揭示并发源、远距离及域差距为主要挑战。
AI中文摘要:
家庭环境中的具身助手必须推断发生了什么、发生在何处以及何时发生,以及如何应对。为此,我们提出了RMS-AQA,一个面向真实世界家庭环境的空间音频问答(SAQA)基准。该基准采用双阶段问答(QA)格式,以全面评估音频语言模型(ALMs)的能力:首先对可听声音事件进行定位,然后基于该定位执行复杂的时空推理。为最大化声学真实性,我们的数据集结合了真实的、真实世界的一阶环境立体声(FOA)录音,以及使用实测房间冲激响应(RIRs)生成的高保真合成数据。此外,我们提供了一个轻量级空间插件,可将FOA格式数据注入冻结的音频语言骨干网络。实验结果表明,主要挑战源于并发声源、远距离以及RIR合成与真实录音之间的模拟到真实域差距。
英文摘要:
Embodied assistants in domestic environments must infer what happened, where and when it occurred, and how to respond. To address this, we introduce RMS-AQA, a spatial audio question answering (SAQA) benchmark for real-world domestic environments. The benchmark features a two-stage question-answering (QA) format to comprehensively assess the ability of audio-language models (ALMs) to first ground audible sound events and subsequently perform complex spatio-temporal reasoning based on that grounding. To maximize acoustic realism, our dataset combines authentic real-world first-order Ambisonics (FOA) recordings with high-fidelity synthetic data generated using measured room impulse responses (RIRs). Furthermore, we provide a lightweight spatial plug-in that injects FOA-format data into frozen audio-language backbones. Experimental results reveal that the primary challenges stem from concurrent sources, far distance, and sim-to-real domain gap between RIR-synthesized and authentic recordings.