arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SEA-LM:面向可穿戴麦克风阵列的具身空间音频理解

SEA-LM: Egocentric Spatial Audio Understanding for Wearable Microphone Arrays

Sonal Kumar, Sinan Hersek, Artem Dementyev, Mengzhen Pan, Ishan Chatterjee, Anurag Kumar, Ramani Duraiswami, Dinesh Manocha, Andrea Colaco

arXiv 2610.05610首次发表:更新:

发表机构

University of Maryland, College Park; Google; Google DeepMind(马里兰大学学院公园分校; 谷歌; 谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大型音频语言模型缺乏空间感知的问题,提出SEA-LM模型,通过FOACODER编码器提取空间音频特征,结合多模态大语言模型与时空加权损失,在声音定位和选择性转录任务上优于基线,并适应多种麦克风阵列配置。

AI 中文摘要

具身、以自我为中心的智能从根本上要求具备在复杂环境中理解空间音频的能力。虽然大型音频语言模型在单声道推理方面表现出色,但它们缺乏空间感知,丢弃了关键的、能够实现声音定位并有助于改善重叠声源分离的空间线索。为解决这一问题,我们提出了SEA-LM,一种空间音频理解模型。首先,我们引入了FOACODER,一种布局灵活的空间音频编码器,通过波束成形从可变数量、可变位置的智能眼镜阵列中编码一阶环境立体声,并在源定位和以自我为中心的语音活动检测目标上进行训练。我们训练了一个多模态大语言模型(MLLM),通过涵盖六项任务的两阶段课程来理解这些空间音频嵌入,这些任务包括在多个说话人和重叠声音的环境中进行声音定位和空间选择性转录。为了防止转录输出主导下一个词元预测损失并压倒方向预测,我们引入了时空加权交叉熵损失。在我们的评估集上,SEA-LM在大多数转录任务中实现了比对比基线更低的方位角和仰角平均绝对误差、更高的时间交并比、更低的外部声源幻觉率和缺失声源率,以及更低的词错误率,同时在1,211种具有4至9个麦克风的智能眼镜阵列配置中保持鲁棒性。

英文摘要

Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentanglement of overlapping sound sources. To address this, we present SEA-LM, a Spatial Audio Understanding model. First, we introduce FOACODER, a layout-flexible spatial audio encoder trained on source localization and ego-centric voice activity detection objectives to encode First Order Ambisonics derived from variable-count, variable-position smart-glasses arrays via beamforming. We train a Multimodal Large Language Model (MLLM) to understand these spatial audio embeddings through a two-stage curriculum spanning six tasks, including sound localization and spatially selective transcription in settings with multiple speakers and overlapping sounds. To prevent the transcription outputs from dominating the next token prediction loss and overwhelming the direction predictions, we introduce a Spatio-temporal Weighted Cross-Entropy Loss. On our evaluation set, SEA-LM achieves lower azimuth and elevation MAE, higher temporal IoU, lower external-source hallucination and missing-source rates, and lower WER on most transcription tasks than compared baselines, while remaining robust across 1,211 smart-glasses array configurations with 4 to 9 microphones.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑