arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAID:用于声音事件定位与检测的语义声学成像检测器

SAID: Semantic Acoustic Imaging Detector for Sound Event Localization and Detection

Runbang Wang, Zining Liang, Yin Cao, Qiuqiang Kong

arXiv 2609.31492首次发表:更新:

发表机构

Nanjing University; The Chinese University of Hong Kong; Institute of Acoustics, Chinese Academy of Sciences; Shun Hing Institute of Advanced Engineering(南京大学; 香港中文大学; 中国科学院声学研究所; 信兴高等工程研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有声源定位仅估计方向而无法描述区域和能量的问题,提出SAID模型,通过预训练编码器并联合预测区域、能量和类别,在DCASE2026任务3上取得最优性能。

AI 中文摘要

在日常生活中,人们会听到周围的语言、脚步声和音乐。我们通常能识别这些声音并判断其来源方向。每个声源都可以显示在单独的声学图上,该图是一幅覆盖水平方向360°和垂直方向180°的矩形图像。该图将声源所占据的方向显示为一个区域,并显示该区域内的声能。类别标签用于标识声音。从音频预测这些带标签的声学图被称为语义声学成像。此类图可帮助机器人感知周围环境,并允许增强现实显示器在真实世界上显示声音区域和类别。现有模型能够识别声音类别并为每个声源估计一个方向。然而,仅一个方向并不能描述声源区域或其能量。声学成像还必须区分邻近方向上的声源,而活动声源的数量及其占据的区域会随时间变化。因此,我们提出了语义声学成像检测器(SAID),它从音频中为每个活动声源预测一个单独的带标签声学图。首先,我们通过无类别标签的跨方向声能估计来预训练SAID的音频编码器Audio2Sph。然后,我们训练完整的SAID模型以同时预测声源区域、能量和类别。我们还开发了一个流程,用于生成模拟录音以进行预训练,并支持在真实录音上进行微调。在官方DCASE2026任务3轨道A评估集上,我们提交的系统以0.1080的宏平均平均精度(Macro mAP)和0.3962的宏皮尔逊相关系数(Macro Pearson r)排名第一。演示和代码可在该https URL处获取。

英文摘要

In daily life, people hear speech, footsteps, and music around them. We can often recognize these sounds and judge where they come from. Each sound source can be shown on a separate acoustic map, a rectangular image covering $360^{\circ}$ horizontally and $180^{\circ}$ vertically. The map shows the directions occupied by the source as a region and the sound energy within that region. A class label identifies the sound. Predicting these labeled acoustic maps from audio is called semantic acoustic imaging. Such maps could help robots perceive their surroundings and allow augmented reality displays to show sound regions and classes over the real world. Existing models can recognize sound classes and estimate a direction for each source. However, a direction alone does not describe the source region or its energy. Acoustic imaging must also distinguish sound sources in nearby directions, while the number of active sources and the regions they occupy can change over time. We therefore propose the Semantic Acoustic Imaging Detector (SAID), which predicts a separate labeled acoustic map for each active source from audio. First, we pretrain Audio2Sph, SAID's audio encoder, through sound energy estimation across directions without class labels. Then, we train the complete SAID model to predict source regions, energy, and classes together. We also develop a pipeline that generates simulated recordings for pretraining and supports fine-tuning on real recordings. On the official DCASE2026 Task 3 Track A evaluation set, our submitted system ranks first with 0.1080 macro-averaged mean average precision (Macro mAP) and 0.3962 Macro Pearson $r$. Demos and code are provided at https://github.com/IN03X/SAID.

Comments5 pages, 3 figures. Accepted at DCASE 2026 Workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑