长语音分离中跨段置换对齐的动态聚类方法
Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation
- Tampere University(坦佩雷大学)
- Nokia Technologies(诺基亚科技)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种兼容现有分离模型的即插即用动态聚类后处理模块,通过说话人嵌入参考池实现跨段置换对齐,在长语音分离的密集、稀疏场景及未知说话人数量场景下性能优于现有方法。
AI中文摘要:
长语音分离通常采用分块-处理-拼接范式,即将录音划分为短片段、独立处理后再拼接,其挑战在于预测跨段置换。本文提出一种无需训练的动态聚类方法,用于基于说话人嵌入参考池的跨段置换对齐,该方法通过当前片段嵌入与参考池的余弦相似度预测置换,并根据与现有参考的整体余弦相似度保留最具代表性的说话人嵌入以更新参考池。作为即插即用的后处理模块,该方法可兼容现有分离模型,在密集和稀疏长语音场景下均表现出优于现有方法的性能,尤其在具有长话语间隙的挑战性稀疏场景中表现突出,且在未知说话人数量场景下对说话人数量估计误差具有鲁棒性。
英文摘要:
Long speech separation typically employs a segment-separation-stitch paradigm where recordings are divided into short segments, processed independently, and stitched together. Its challenge lies in predicting cross-segment permutations. This paper proposes a training-free dynamic clustering approach for cross-segment permutation alignment using speaker embedding reference pools. The method predicts the permutation using the cosine similarity between current segment embeddings and the reference pools. The approach updates reference pools by retaining the most representative speaker embeddings based on their overall cosine similarity with existing references. As a plug-and-play post-processing module compatible with existing separation models, the proposed method demonstrates superior performance compared to existing methods on dense and sparse long speech scenarios, particularly in challenging sparse scenarios with extended utterance gaps, and further shows robustness to speaker count estimation errors in unknown speaker count scenarios.