发表机构
Ohio State University; Meta Reality Labs; The Chinese University of Hong Kong, Shenzhen(俄亥俄州立大学; Meta现实实验室; 香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对说话人跟踪中重叠语音和说话人数变化问题,提出分段在线模块化框架,采用分离、VAD及两阶段组织策略,在LibriCSS和AMI上达到最优性能,显著降低DER和cpWER。
AI 中文摘要
说话人跟踪是在时间上分离并跟踪多个说话人的任务。它必须处理重叠语音、语音起始和结束、说话人身份以及随时间变化的说话人数。我们提出了一种用于单通道和多通道说话人跟踪的分段在线模块化框架。所提出的系统首先在每个分段中执行说话人分离,并通过语音活动检测(VAD)计算语音活动。为了随时间生成说话人轨迹,我们引入了一种两阶段顺序组织策略:基于重叠的拼接用于连续分组,基于记忆的说话人验证用于不连续分组。对于说话人分离,我们采用复数谱映射来估计底层说话人的实部和虚部谱图。所提出的系统在LibriCSS和AMI数据集上实现了最先进的分段在线跟踪性能。与其他方法相比,我们的框架显著降低了说话人日志错误率(DER)和拼接最小置换词错误率(cpWER)。
英文摘要
Speaker tracking is the task of separating and following multiple speakers over time. It must address overlapped speech, speech onset and offset, talker identity, and time-varying speaker count. We propose a segment-online, modular framework for single- and multi-channel speaker tracking. The proposed system first performs speaker separation in each segment and computes speech activity via voice activity detection (VAD). To generate speaker tracks over time, we introduce a two-stage sequential organization strategy: Overlap-based stitching for continuous grouping and memory-based speaker verification for discontinuous grouping. For speaker separation, we employ complex spectral mapping to estimate the real and imaginary spectrograms of underlying speakers. The proposed system achieves state-of-the-art segment-online tracking performance on the LibriCSS and AMI datasets. Our framework significantly reduces diarization error rate (DER) and concatenated minimum-permutation word error rate (cpWER) compared to other methods.