arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于深度神经网络的声学混合信号中语音、音乐与噪声功率的频率依赖估计,用于助听器场景分析

DNN-Based Frequency-Dependent Estimation of Speech, Music, and Noise Power in Acoustic Mixtures for Hearing-Aid Scene Analysis

Mats Lang, Thomas Haubner, Nina Kiessling, Christoph Hoog Antink, Henning Puder

arXiv 2608.17482首次发表:更新:

AI 中文总结

针对助听器场景分析中多独立估计器的问题,提出用深度神经网络估计混合信号的语音、音乐、噪声功率分量,以VAD验证其性能与SOTA相当且场景描述更丰富。

AI 中文摘要

声学场景分析对于将助听器信号处理算法适配到当前聆听环境至关重要。然而,最先进(SOTA)系统通常依赖多个独立估计器完成场景分类、语音活动检测(VAD)或信噪比估计等任务,这会增加计算复杂度,且无法利用相关任务间的依赖关系。为解决该问题,我们受助听器用户典型聆听目标的启发,提出一种统一且可解释的声学场景表示,将观测混合信号频谱分解为语音、音乐与噪声功率分量。具体而言,我们通过因果低复杂度深度神经网络估计时间与频率依赖的功率占比,理论上可通过简单后处理衍生出多个下游声学场景分析指标。本研究以VAD为代表性下游任务验证该表示,结果显示其性能与SOTA估计器相当,同时提供了丰富得多的场景描述。

英文摘要

Acoustic scene analysis is essential for adapting hearing-aid signal processing algorithms to the current listening environment. However, state-of-the-art (SOTA) systems typically rely on multiple independent estimators for tasks such as scene classification, Voice Activity Detection (VAD), or Signal-to-Noise Ratio estimation, which increases computational complexity and fails to exploit dependencies between related tasks. To address this problem, we propose a unified and interpretable acoustic scene representation by decomposing the observed mixture spectrum into speech, music, and noise power components. This is motivated by the typical listening targets of hearing-aid users. In particular, we estimate time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing. In this work, we validate the proposed representation using VAD as a representative downstream task and show performance comparable to a SOTA estimator while providing a substantially richer scene description.

CommentsAccepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑