研究基于基础模型的说话人聚类中空间信息的整合
Investigating the Integration of Spatial Information in Foundation-Model-Based Speaker Diarization
浏览论文内容
中文总结 AI 辅助
研究基于基础模型的说话人聚类中空间信息整合,比较波束形成器与基础模型级联、多通道基础模型、基于提取空间特征对下游网络条件设定三种方法,发现条件设定法性能最佳且能消除相关错误,是有竞争力的方法。
中文摘要 AI 辅助
从多通道输入中获取的空间信息已被证明能改善诸如聚类和源分离等会议处理任务。同时,基于大型预训练单通道基础模型(如WavLM)提取的特征进行的聚类取得了当前最优性能。本文比较了三种将空间特征整合到基于基础模型的聚类系统中的方法:波束形成器与单通道基础模型的级联、多通道基础模型以及基于明确提取空间特征对下游网络进行条件设定。结果表明,波束形成器前端在语音重叠区域甚至对聚类性能有害,而条件设定方法性能最佳,这表明纳入明确的空间特征是基础模型支持的聚类的一种有竞争力的方法。该方法还进行了详细的错误分析,表明条件设定系统在很大程度上消除了仅使用频谱或仅使用空间特征时会出现的错误。
英文摘要
Spatial information gleaned from multi-channel input has been shown to lead to improvements in meeting processing tasks like diarization and source separation. At the same time, diarization based on features extracted by large pretrained single-channel foundation models, such as WavLM, achieved state-of-the-art performance. This work compares three approaches to integrate spatial features into foundation model-based diarization systems: the cascade of a beamformer and a single-channel foundation model, a multi-channel foundation model, and the conditioning of the downstream network on explicitly extracted spatial features. Results show that the beamformer front-end is even detrimental to diarization performance in regions of overlapped speech, while best performance is achieved with the conditioning, demonstrating that the incorporation of explicit spatial features is a competitive approach to foundation-model-supported diarization. This approach is further subjected to a detailed error analysis showing that the conditioning system removes errors to a good extent that would occur when either only spectral or only spatial features were used.