arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MiDashengLM-Spatial:统一通用音频理解与空间感知

MiDashengLM-Spatial: Unifying General Audio Understanding and Spatial Awareness

Jinbo Hu, Hang Su, Lichun Fan, Heinrich Dinkel, Gang Li, Zhanchen Dai, Yiru Zhang, Chang Liu, Peng Wang, Junnan Wu, Jian Luan, Cong Zou, Heng Qu

arXiv 2610.11156首次发表:更新:

发表机构

Xiaomi Inc.(小米公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MiDashengLM-Spatial是首个开源端到端统一音频-语言模型,通过新增Spatial-Dasheng空间音频编码器,在不损害通用音频理解的前提下实现了优异的空间感知能力,在相关基准上表现突出。

AI 中文摘要

大型音频-语言模型(LALMs)在通用音频理解方面已取得优异性能,但多数模型针对单声道输入设计,会丢弃空间感知必需的通道间线索。相比之下,现有空间音频-语言模型专为空间任务构建,无法利用单声道LALMs的通用理解能力。我们提出MiDashengLM-Spatial,据我们所知,这是首个开源端到端统一音频-语言模型,可在单一架构内同时支持通用音频理解与空间感知。该模型通过分层语义到空间条件模块,将空间音频编码器Spatial-Dasheng扩展到MiDashengLM中,该模块在多个深度将中间语义表示注入空间分支,同时保留原始语义通路。为大规模提供空间音频-语言监督,我们开发了数据合成流水线,可渲染带有场景级空间描述和问答对的多样空间声学场景。实验表明,Spatial-Dasheng在真实场景的声音事件定位与检测任务中表现优异,MiDashengLM-Spatial在空间理解与推理基准上显著优于现有LALMs,同时在多样单声道基准上与最先进的8B规模LALMs保持竞争力,证明可在不损害通用音频理解的前提下获取空间感知能力。源代码和模型检查点可在两个指定URL获取。

英文摘要

Large audio-language models (LALMs) have achieved strong performance in general audio understanding, yet most are designed for monaural input and discard the inter-channel cues essential for spatial perception. In contrast, existing spatial audio-language models are purpose-built for spatial tasks and fail to capitalize on the general understanding capabilities of monaural LALMs. We present MiDashengLM-Spatial, the first open-source end-to-end unified audio-language model, to our knowledge, which supports both general audio understanding and spatial awareness within a single architecture. It extends MiDashengLM with a spatial audio encoder, Spatial-Dasheng, integrated through a hierarchical semantic-to-spatial conditioning module that injects intermediate semantic representations into the spatial branch at multiple depths while preserving the original semantic pathway. To provide spatial audio-language supervision at scale, we develop a data synthesis pipeline that renders diverse spatial acoustic scenes with scene-level spatial descriptions and question-answer pairs. Experiments show that Spatial-Dasheng achieves strong performance on sound event localization and detection in real-world scenes, and that MiDashengLM-Spatial substantially outperforms existing LALMs on spatial understanding and reasoning benchmarks. Meanwhile, it remains competitive with state-of-the-art 8B-scale LALMs on diverse monaural benchmarks, demonstrating that spatial awareness can be acquired without compromising general audio understanding. The source code and model checkpoint are available at https://github.com/xiaomi-research/midashenglm-spatial and https://huggingface.co/mispeech/midashenglm-spatial.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑