基于最大加权似然估计的独立低秩矩阵分析与秩约束空间协方差矩阵估计的在线算法
Online Algorithms for Independent Low-Rank Matrix Analysis and Rank-Constrained Spatial Covariance Matrix Estimation Based on Maximum Weighted Likelihood Estimation
浏览论文内容
中文总结 AI 辅助
本文提出基于最大加权似然估计的ILRMA和RCSCME在线算法,通过逐帧代价函数与辅助函数技术推导更新规则,并采用近似与加速技术,实现动态场景下实时多通道语音提取,性能优于传统方法。
中文摘要 AI 辅助
在漫射噪声条件下进行实时多通道语音提取(MSE)是一项重要任务,具有广泛的应用,例如语音识别和助听器。在本文中,我们提出了独立低秩矩阵分析(ILRMA)和秩约束空间协方差矩阵估计(RCSCME)的在线算法。此前,我们提出了基于RCSCME的方法的实时扩展:一种基于ILRMA和RCSCME并使用分块批量算法的MSE方法。然而,它假设空间特性在单个批次内是平稳的,因此在目标说话人移动的动态情况下,其性能可能会下降。为了解决这个问题,我们通过以下三个步骤推导了ILRMA和RCSCME的在线算法。首先,我们基于最大加权似然估计为ILRMA和RCSCME制定了逐帧代价函数。其次,我们基于辅助函数技术推导了逐帧代价函数的更新规则。这些朴素的更新规则在实际机器上实时执行时计算成本高昂。因此,我们最终通过用其估计值近似某些中间参数来推导在线算法。此外,我们为这些在线算法提出了稳定化和进一步加速的技术。在实验中,我们模拟了目标说话人静止或移动的情况,并表明所提出的方法相比传统方法实现了优越的语音提取性能。此外,使用真实世界记录的信号,我们证明了所提出方法在实际场景中的有效性。
英文摘要
Real-time multichannel speech extraction (MSE) under diffuse noise conditions is an important task with a wide range of applications, such as speech recognition and hearing aids. In this paper, we propose online algorithms for independent low-rank matrix analysis (ILRMA) and rank-constrained spatial covariance matrix estimation (RCSCME). Previously, we proposed a real-time extension of the RCSCME-based method: an MSE method based on ILRMA and RCSCME using the blockwise batch algorithm. However, it assumes that the spatial characteristics are stationary within a single batch, and thus, in dynamic situations where the target speaker moves, its performance may degrade. To address this problem, we derive the online algorithms for ILRMA and RCSCME in the following three steps. First, we formulate framewise cost functions for ILRMA and RCSCME on the basis of maximum weighted likelihood estimation. Second, we derive the update rules for the framewise cost functions on the basis of auxiliary-function techniques. These naive update rules are computationally costly for real-time execution on a practical machine. Thus, we finally derive the online algorithms by approximating some intermediate parameters with their estimates. Furthermore, we propose stabilization and further acceleration techniques for these online algorithms. In experiments, we simulate situations where a target speaker is stationary or moves and show that the proposed method achieves superior speech extraction performance compared with conventional methods. In addition, using real-world recorded signals, we demonstrate the effectiveness of the proposed method in practical scenarios.
发表机构
- The University of Tokyo(东京大学)
- National Institute of Technology, Kagawa College(国立工业高等专门学校香川高等专门学校)
- Yamaha Corporation(雅马哈公司)
机构由 AI 辅助整理,请以论文原文为准。