发表机构
Xiaomi Inc.(小米公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MiDashengLM-Spatial是首个开源端到端统一音频-语言模型,通过新增Spatial-Dasheng空间音频编码器,在不损害通用音频理解的前提下实现了优异的空间感知能力,在相关基准上表现突出。
AI 中文摘要
大型音频-语言模型(LALMs)在通用音频理解方面已取得优异性能,但多数模型针对单声道输入设计,会丢弃空间感知必需的通道间线索。相比之下,现有空间音频-语言模型专为空间任务构建,无法利用单声道LALMs的通用理解能力。我们提出MiDashengLM-Spatial,据我们所知,这是首个开源端到端统一音频-语言模型,可在单一架构内同时支持通用音频理解与空间感知。该模型通过分层语义到空间条件模块,将空间音频编码器Spatial-Dasheng扩展到MiDashengLM中,该模块在多个深度将中间语义表示注入空间分支,同时保留原始语义通路。为大规模提供空间音频-语言监督,我们开发了数据合成流水线,可渲染带有场景级空间描述和问答对的多样空间声学场景。实验表明,Spatial-Dasheng在真实场景的声音事件定位与检测任务中表现优异,MiDashengLM-Spatial在空间理解与推理基准上显著优于现有LALMs,同时在多样单声道基准上与最先进的8B规模LALMs保持竞争力,证明可在不损害通用音频理解的前提下获取空间感知能力。源代码和模型检查点可在两个指定URL获取。
英文摘要
Large audio-language models (LALMs) have achieved strong performance in general audio understanding, yet most are designed for monaural input and discard the inter-channel cues essential for spatial perception. In contrast, existing spatial audio-language models are purpose-built for spatial tasks and fail to capitalize on the general understanding capabilities of monaural LALMs. We present MiDashengLM-Spatial, the first open-source end-to-end unified audio-language model, to our knowledge, which supports both general audio understanding and spatial awareness within a single architecture. It extends MiDashengLM with a spatial audio encoder, Spatial-Dasheng, integrated through a hierarchical semantic-to-spatial conditioning module that injects intermediate semantic representations into the spatial branch at multiple depths while preserving the original semantic pathway. To provide spatial audio-language supervision at scale, we develop a data synthesis pipeline that renders diverse spatial acoustic scenes with scene-level spatial descriptions and question-answer pairs. Experiments show that Spatial-Dasheng achieves strong performance on sound event localization and detection in real-world scenes, and that MiDashengLM-Spatial substantially outperforms existing LALMs on spatial understanding and reasoning benchmarks. Meanwhile, it remains competitive with state-of-the-art 8B-scale LALMs on diverse monaural benchmarks, demonstrating that spatial awareness can be acquired without compromising general audio understanding. The source code and model checkpoint are available at https://github.com/xiaomi-research/midashenglm-spatial and https://huggingface.co/mispeech/midashenglm-spatial.