发表机构
New Jersey Institute of Technology; Mitsubishi Electric Research Laboratories (MERL)(新泽西理工学院; 三菱电机研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估推理时算法以改进多阶段声音场景语义分割中的检测与提取,验证了多种设计选择在DCASE 2025任务4数据集上的效果,并为现实场景应用奠定基础。
AI 中文摘要
声音场景的空间语义分割(S5)涉及在音频文件中检测和提取目标声音事件。然而,在典型的检测-提取流程中,检测模块的错误会传播到提取模块,从而降低整体性能。在本工作中,我们开发了一个多通道检测-提取模型,并评估了推理时算法以改进多阶段设置中的类别检测和目标声音提取。我们评估了两种适应度函数:二元交叉熵和混合一致性;两种重新标记策略:朴素和条件混淆矩阵重新标记;以及三种源估计评估模型:多标签分类器、微调的单源分类器和现成的音频评判器。我们彻底评估了这些设计选择对DCASE 2025任务4数据集的影响,并在留出评估集上验证了基于熵的推理时更新门控,为S5系统在现实场景中的推理时更新奠定了基础。
英文摘要
Spatial semantic segmentation of sound scenes (S5) involves the detection and extraction of target sound events in audio files. However, in typical detection-extraction pipelines, errors in the detection module propagate to the extraction module, degrading overall performance. In this work, we develop a multichannel detection-extraction model and evaluate inference-time algorithms to improve the class detection and target sound extraction in a multi-stage setup. We evaluate two fitness functions: binary cross entropy and mixture consistency; two relabeling strategies: naive and conditional confusion matrix relabeling; and three source estimate evaluation models: a multi-label classifier, a fine-tuned single-source classifier, and an off-the-shelf audio judge. We thoroughly evaluate the impact of these design choices on the DCASE 2025 Task 4 dataset, and validate entropy-based gating of inference-time updates on the held-out evaluation set, laying the groundwork for inference-time updates of S5 systems in real-world scenarios.
CommentsAccepted to DCASE 2026