arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

过紧化感知的伪标签方法用于紧边界说话人日志

Over-Tightening-Aware Pseudo-Labeling for Tight-Boundary Speaker Diarization

Shota Horiguchi, Takanori Ashihara, Marc Delcroix, Naohiro Tawara, Alexis Plaquet

arXiv 2609.09965首次发表:更新:

发表机构

NTT, Inc.(日本电报电话公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对说话人日志中伪标签过紧化导致漏检的问题,提出三种改进方法,减少漏检并提升日志及下游多说话人ASR性能。

AI 中文摘要

在松散标签(如带有填充边界或填充停顿的语音片段)上训练说话人日志模型,通常会导致模型输出同样松散。为了获得更紧的边界,已有研究提出基于因果和反因果模型平均输出的伪标签方法。然而,由于伪标签是基于估计的,它们可能遭受过紧化问题,这会增加漏检,并可能作为不可恢复的错误传播到下游任务。本文仔细分析了过紧化的原因,并提出了三种应对方法:(i) 移除停顿填充而非填充边界,(ii) 引入预热阶段以减轻因果和反因果预测起始附近的漏检,(iii) 使基于伪标签的协同训练感知用于最终推理的非因果模型。实验结果表明,所提方法减少了由过紧化引起的漏检,并提高了日志准确性和下游多说话人自动语音识别性能。

英文摘要

Training speaker diarization models on loose labels, such as speech segments with padded boundaries or filled pauses, often results in similarly loose model outputs. To obtain tighter boundaries, pseudo-labeling based on the averaged outputs of causal and anticausal models has been proposed. However, since the pseudo-labels are estimation-based, they can suffer from over-tightening, which increases missed detections that can propagate as unrecoverable errors to downstream tasks. This paper carefully analyzes the causes of over-tightening and proposes three approaches to address them: (i) removing pause filling rather than padding, (ii) introducing a burn-in phase to mitigate missed detections near the beginning of causal and anticausal predictions, and (iii) making pseudo-label-based co-training aware of the non-causal model used for final inference. Experimental results show that the proposed method reduces missed detections caused by over-tightening and improves both diarization accuracy and downstream multi-talker ASR performance.

CommentsAccepted to IEEE SLT 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑