arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAIL:基于大语言模型的空间音频智能,通过解耦声学-空间编码与双流Q-Former实现

SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former

Zhengding Luo, Jinyang Wu, Haozhe Ma, Yanghao Zhou, Woon-Seng Gan, Wenwu Wang

arXiv 2609.34347首次发表:更新:

发表机构

Nanyang Technological University; Singapore Management University; Tencent Hy Frontier Lab; Beijing Institute of Technology; University of Surrey(南洋理工大学; 新加坡管理大学; 腾讯Hy前沿实验室; 北京理工大学; 萨里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SAIL提出解耦声学-空间编码与双流Q-Former,保持源级对应,提升多声源空间理解与推理性能。

AI 中文摘要

空间音频大语言模型使具身智能体、可穿戴助手和沉浸式系统能够识别声音事件、定位声源并推理其空间关系。然而,现有的空间音频大语言模型通常依赖声学与空间特征的早期融合以及源无关的令牌表示。这些设计难以保持单个声音事件与其空间属性之间的对应关系,尤其是在多声源场景中。为解决这一局限,我们提出了SAIL,一种基于大语言模型的空间音频智能框架,从音频编码到LLM对齐过程中保持声学-空间结构和源级对应关系。SAIL引入了解耦空间音频变换器,将梅尔频谱图和耳间相位差特征表示为独立的声学流和空间流。源判别任务查询进一步学习每个声源的事件、方向和距离信息。双流Q-Former随后利用按声源槽组织的声学查询和空间查询,将两个流与LLM对齐。与早期融合基线相比,SAIL在双声源声音事件检测、方向和距离估计以及空间推理方面均取得了一致的改进。这些结果证明了结构化、源判别的音频表示对于多声源空间理解和推理的重要性。

英文摘要

Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑