arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于嘈杂和混响环境中语音增强的房间冲激响应嵌入

Room Impulse Response Embeddings for Speech Enhancement in Noisy and Reverberant Environments

Adrian Meise, Reinhold Haeb-Umbach

arXiv 2609.31041首次发表:更新:

发表机构

Paderborn University(帕德博恩大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出自监督方法学习房间冲激响应嵌入,用于在混响和噪声环境中提升语音增强性能,并在所有指标上取得一致改进。

AI 中文摘要

我们提出了一种自监督方法,用于从单通道带噪混响语音中学习房间冲激响应(RIR)表示。该方法首先在混响数据上训练,然后在带噪混响数据上训练,最后采用教师-学生方法,其中学生模型在给定混响输入的带噪版本时,学习复制教师模型的嵌入。我们通过从这些嵌入中估计声学房间参数来评估其表示能力。将判别式语音增强模型基于这些嵌入进行条件化处理,在所有评估指标(包括下游词错误率)上均取得了一致的性能提升,对混响和带噪混响语音均有效。

英文摘要

We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student approach, where the student learns to replicate the teacher's embeddings when given a noisy version of the reverberant input. We assess their representational capabilities by estimating acoustic room parameters from them. Conditioning a discriminative speech enhancement model on the derived embeddings yields consistent gains across all evaluated metrics, including downstream word error rate, for both reverberant and noisy-reverberant speech.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑