发表机构
LTCI, Telecom Paris, Institut Polytechnique de Paris; LIUM, Universite du Mans(LTCI、巴黎电信学院、巴黎综合理工学院; LIUM、勒芒大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对神经 FCASA 方法提出带 Beta 说话人活动先验的贝叶斯 diarization 模型,在 AMI 数据集上使 diarization 和 Jaccard 错误率显著降低,提升了远场说话人 diarization 的鲁棒性。
AI 中文摘要
远场说话人 diarization 因恶劣声学环境、说话人数量变化及语音重叠而极具挑战性。尽管数据驱动方法已展现出色性能,模型驱动方法通过利用多通道录音的空间信息提供了极具吸引力的替代方案。本文旨在为名为神经 FCASA 的模型驱动方法提出一种贝叶斯 diarization 模型以增强其鲁棒性。具体而言,我们针对说话人活动提出 Beta 先验,进而提出一种变分下界目标,该目标可视为正则化连续说话人活动分数,用于替代原始交叉熵损失以训练 diarization 模型。实验结果表明,在 AMI 数据集上,与基线相比,diarization 错误率至少降低 3%(相对降低 16%),Jaccard 错误率至少降低 4%(相对降低 20%),取得显著提升。
英文摘要
Distant speaker diarization remains challenging due to adverse acoustic conditions, varying numbers of speakers and overlapping speech. While data-driven approaches have shown strong performance, model-driven methods offer a compelling alternative by leveraging spatial information from multichannel recordings. This paper is motivated to propose a Bayesian diarization model for a model-driven method called neural FCASA to enhance its robustness. Specifically, we propose a beta prior over speaker activity and hence a variational lower bound objective that can be seen as a regularized continuous speaker activity score in place of the original cross-entropy loss to train the diarization model. Our experiments show significant improvements in terms of Diarization Error Rate by at least 3% (16% relatively) and Jaccard Error Rate by at least 4% (20% relatively) on the AMI dataset compared to the baseline.
CommentsAccepted by Interspeech 2026