AI 中文总结
针对助听器场景分析中多独立估计器的问题,提出用深度神经网络估计混合信号的语音、音乐、噪声功率分量,以VAD验证其性能与SOTA相当且场景描述更丰富。
AI 中文摘要
声学场景分析对于将助听器信号处理算法适配到当前聆听环境至关重要。然而,最先进(SOTA)系统通常依赖多个独立估计器完成场景分类、语音活动检测(VAD)或信噪比估计等任务,这会增加计算复杂度,且无法利用相关任务间的依赖关系。为解决该问题,我们受助听器用户典型聆听目标的启发,提出一种统一且可解释的声学场景表示,将观测混合信号频谱分解为语音、音乐与噪声功率分量。具体而言,我们通过因果低复杂度深度神经网络估计时间与频率依赖的功率占比,理论上可通过简单后处理衍生出多个下游声学场景分析指标。本研究以VAD为代表性下游任务验证该表示,结果显示其性能与SOTA估计器相当,同时提供了丰富得多的场景描述。
英文摘要
Acoustic scene analysis is essential for adapting hearing-aid signal processing algorithms to the current listening environment. However, state-of-the-art (SOTA) systems typically rely on multiple independent estimators for tasks such as scene classification, Voice Activity Detection (VAD), or Signal-to-Noise Ratio estimation, which increases computational complexity and fails to exploit dependencies between related tasks. To address this problem, we propose a unified and interpretable acoustic scene representation by decomposing the observed mixture spectrum into speech, music, and noise power components. This is motivated by the typical listening targets of hearing-aid users. In particular, we estimate time- and frequency-dependent power proportions by a causal low-complexity Deep Neural Network, from which multiple downstream acoustic scene analysis measures can in principle be derived by simple post-processing. In this work, we validate the proposed representation using VAD as a representative downstream task and show performance comparable to a SOTA estimator while providing a substantially richer scene description.
CommentsAccepted to International Workshop on Acoustic Signal Enhancement (IWAENC) 2026