发表机构
Ghent University(根特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对流式语音增强提出两种频谱加权STFT损失,设计HyST-Net骨干网络,实验表明所提损失可改善高频重建,频谱自适应损失还能增强中频,实现更均衡的全频段频谱重建。
AI 中文摘要
本文提出两种频谱加权短时傅里叶变换(STFT)损失函数,用于轻量流式语音增强,以解决由幅度-相位补偿效应导致的中高频区域幅度过度衰减问题。所提出的sigmoid加权损失对相位感知贡献应用平滑的频率相关调制,而依赖信号的频谱自适应损失则进一步将调制条件设置为真实对数幅度语谱图。为评估所提出的目标,我们额外设计了HyST-Net,这是一种轻量且具有竞争力的骨干网络,采用混合多头注意力门控循环单元(MHA-GRU)频谱-时间建模,适用于低延迟流式场景。实验结果表明,两种损失在高频频谱重建方面均实现了一致的改进;频谱自适应损失进一步增强了中频区域,从而在全频率范围内实现更均衡的频谱重建。
英文摘要
This paper proposes two spectrally weighted STFT loss functions for lightweight streaming speech enhancement, addressing the magnitude over-attenuation in mid-to-high frequency regions caused by the magnitude-phase compensation effect. The proposed sigmoid-weighted loss applies a smooth frequency-dependent modulation to the phase-aware contribution, while the signal-dependent spectrally adaptive loss further conditions the modulation on the ground-truth log-magnitude spectrogram. To evaluate the proposed objectives, we additionally design HyST-Net, a lightweight and competitive backbone with hybrid MHA-GRU spectral-temporal modelling for low-latency streaming scenarios. Experimental results exhibit consistent improvements in high-frequency spectral reconstruction for both losses. The spectrally adaptive loss further enhances the mid-frequency region, resulting in a more balanced spectral reconstruction across the full frequency range.
CommentsAccepted by IWAENC 2026