arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11627eess.AScs.AIeess.SP

基于深度学习的多声源与多麦克风相对传递矩阵估计

Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones

Oshan A. B. Yalegama, Wageesha N. Manamperi

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对多声源多麦克风场景,提出三种基于深度学习的相对传递矩阵估计框架,实验表明其估计精度优于协方差方法,语音增强性能与基线相当。

中文摘要 AI 辅助

相对传递矩阵(ReTM)是近期提出的、适用于多接收端与多声源的相对传递函数的泛化形式,在噪声环境下的语音增强任务中展现出良好性能。利用多通道记录的协方差矩阵估计声源的ReTM对实际应用极具价值,且是迄今为止唯一被提出的相关方法。本文研究基于深度学习的ReTM估计,提出三种新颖的监督学习框架,分别采用时域和短时傅里叶变换域的卷积网络,以及基于长短期记忆(LSTM)的循环神经网络。实验结果表明,所提模型在五项客观指标上,相比基于协方差的方法能实现更准确的ReTM估计;同时验证了所提框架在语音增强任务中的有效性,其性能与基线方法相当。

英文摘要

The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and sources, shows promising performance when applied to speech enhancement in noisy environments. Estimating the ReTM of sound sources by exploiting the covariance matrices of multichannel recordings is highly beneficial for practical applications and, to date, remains the only proposed approach. This paper investigates deep learning-based ReTM estimation. We propose three novel supervised learning frameworks using time and short-time frequency transform domain convolutional networks, and a Long Short-Term Memory-based recurrent neural network. Experimental results demonstrate that the proposed models achieve more accurate estimation of the ReTM using five objective metrics compared to the covariance-based method. We also show the effectiveness of the proposed frameworks for speech enhancement, achieving performance on par with the baseline method.

发表机构

  • University of Moratuwa(莫拉图瓦大学)
  • The Australian National University(澳大利亚国立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑