神经多通道远场说话人日志与重尾源分离模型
Neural Multichannel Distant Speaker Diarization with Heavy-tailed Source Separation Model
- LTCI, Telecom Paris, Institut Polytechnique Paris(巴黎综合理工学院电信巴黎分校LTCI)
- SJTU Paris Elite Institute of Technology, Shanghai Jiao Tong University(上海交通大学巴黎卓越工程师学院)
- LIUM, Le Mans University(勒芒大学LIUM)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种结合重尾分布(尖峰广义高斯和学生t分布)的神经FCASA模型,用于远场多通道说话人日志,在DER和JER指标上较基线取得大幅改进。
AI中文摘要:
远场说话人日志由于困难的声学环境、说话人数目变化和重叠语音而仍然具有挑战性。模型驱动方法被提出以利用多通道录音中有助于日志的语音源特征。本文推广了一个神经模型,该模型联合学习在语音混合物上执行盲源分离和日志(神经FCASA),并采用重尾模型。原始源分离模型中广泛使用高斯分布进行方差建模,我们将其替换为两类重尾模型(尖峰广义高斯分布和学生t分布),以更好地捕获语音信号中的重尾性。得益于高斯尺度混合模型,我们能够将所提方法和原始方法统一在相同形式的学习目标下。我们的实验表明,在各种语料库上,与基线相比,日志错误率(DER)和Jaccard错误率(JER)均获得一致的大幅改进。
英文摘要:
Distant speaker diarization remains challenging due to difficult acoustic environments, varying numbers of speakers and overlapping speech. Model-driven methods are proposed to exploit the speech source features in multi-channel recordings that help diarization. This paper generalizes a neural model that jointly learns to perform blind source separation and diarization over speech mixtures (neural FCASA) with heavy-tailed models. The popular Gaussian distribution has been applied for variance modeling in the original source separation model, which we replace with two families of heavy-tailed models (Leptokurtic Generalized Gaussian distribution and Student's t distribution) to better capture the heavy-tailedness in speech signals. Thanks to the Gaussian scale mixture model, we are able to unify the proposed method and the original one under the same form of learning objective. Our experiments show consistent large improvements in Diarization Error Rate (DER) and Jaccard Error Rate (JER) compared to the baseline on various corpora.