arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

采用精确房间声学模拟训练DeepFilterNet可提升单通道语音增强效果

Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement

Alessia Milo, Georg Götz, Steinar Guðjónsson, Daniel Gert Nielsen, Jesper Pedersen, Finnur Pind

arXiv 2608.20971首次发表:更新:

发表机构

Treble Technologies(特雷布尔科技公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究对比不同保真度的RIR数据集训练DeepFilterNet3,发现采用更高保真度数据集可提升模型的客观指标与ASR性能,增强其对实测环境的泛化能力。

AI 中文摘要

本研究探讨合成房间冲激响应(RIR)数据集的逼真度对用于单通道语音增强的DeepFilterNet3训练的影响。我们将DNS4图像源法(ISM)RIR数据集与采用混合基于波和几何声学模拟生成的更高声学保真度数据集进行对比。我们未单独分析各模拟因素,而是在保持增强模型不变的情况下对比完整的RIR生成流程。通过客观语音增强指标及下游自动语音识别(ASR)对未见过的实测RIR上的模型进行评估。与ISM数据集相比,采用更高保真度数据集训练的模型在客观指标上持续取得小幅提升,且ASR词错误率显著降低。尽管实验未将这些增益归因于单个建模组件,但结果表明,提高合成声学训练数据的整体逼真度可增强DeepFilterNet3对未见过的实测环境的泛化能力。

英文摘要

We investigate how the realism of synthetic room impulse response (RIR) datasets affects the training of DeepFilterNet3 for single-channel speech enhancement. We compare a DNS4 image-source-method (ISM) RIR dataset with a higher-acoustic-fidelity dataset generated using hybrid wave-based and geometrical acoustics simulation. Rather than isolating individual simulation factors, we compare complete RIR generation pipelines while keeping the enhancement model unchanged. Models are evaluated on unseen measured RIRs using objective speech enhancement metrics and downstream automatic speech recognition (ASR). Training with the higher-fidelity dataset consistently yields modest improvements in objective metrics and substantially lower ASR word error rates than the ISM dataset. Although the experiments do not attribute these gains to individual modelling components, they show that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.

Comments5 pages, 2 figures, IWAENC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑