发表机构
KU Leuven(荷语鲁汶大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对选择性听觉注意解码中短窗口决策不可靠的问题,提出多尺度高斯混合模型联合估计发射参数,提高HMM后处理精度,尤其在数据受限时优势明显。
AI 中文摘要
选择性听觉注意解码(sAAD)从脑电图(EEG)中推断在多说话人场景中听者关注的是哪位说话人。一种广泛使用的方法从EEG重建所关注的语音包络,将其与候选包络进行相关分析,并选择皮尔逊相关系数最高的说话人。然而,每个决策窗口中的相关性具有较高的固有方差,使得逐窗口决策不可靠,尤其是在短窗口情况下。近期工作引入了一种隐马尔可夫模型(HMM),该模型随时间整合相关性证据,并需要依赖于状态的发射分布。由于在实践中标记的注意状态通常不可用,这些分布可以通过拟合高斯混合模型(GMM)从未标记的观测中推断出来。然而,当底层高斯分布接近时,在有限样本下无监督拟合变得不太可靠。因此,我们提出了一种多尺度GMM,其参数在多个窗口长度上联合估计。该模型利用了Fisher变换后相关性估计的窗口长度依赖性,其均值近似稳定,而方差随每个窗口的样本数减少而降低。实验表明,多尺度GMM比单尺度拟合更准确地估计发射参数,尤其是在短窗口情况下。其HMM后处理的准确性优势在短录音中最为显著,并随着更多数据的可用而缩小,这促使在数据受限的设置中采用多尺度拟合。
英文摘要
Selective auditory attention decoding (sAAD) infers from electroencephalography (EEG) which speaker a listener attends to in multi-speaker scenarios. A widely used approach reconstructs the attended speech envelope from EEG, correlates it with candidate envelopes, and selects the speaker with the highest Pearson correlation. However, correlations in each decision window have high intrinsic variance, making per-window decisions unreliable, especially for short windows. Recent work introduced a hidden Markov model (HMM) that integrates correlation evidence over time and requires state-dependent emission distributions. Since labeled attention states are typically unavailable in practice, these distributions can be inferred from unlabeled observations by fitting a Gaussian mixture model (GMM). However, when the underlying Gaussians are close, unsupervised fitting becomes less reliable with limited samples. We therefore propose a multiscale GMM whose parameters are jointly estimated across multiple window lengths. The model exploits the window-length dependence of Fisher-transformed correlation estimates, whose means are approximately stable while their variances decrease with the number of samples per window. Experiments show that the multiscale GMM estimates emission parameters more accurately than single-scale fitting, particularly at short windows. Its HMM-post-processed accuracy advantage is largest for short recordings and narrows as more data become available, motivating multiscale fitting in data-constrained settings.
Comments5 pages, 3 figures, conference paper (submitted to ICASSP)