超越DER:拥挤场景下端到端说话人日志中的说话人计数
Beyond DER: Speaker Counting in Crowded End-to-End Diarization
浏览论文内容
中文总结 AI 辅助
针对端到端说话人日志在拥挤场景下低估说话人数的问题,提出门控损失耦合说话人存在性与帧活动性,采用说话人加权正则化,显著降低JER和DER。
中文摘要 AI 辅助
说话人日志必须解决两个问题:统计给定对话中出现的说话人数量,以及将语音分配给每个说话人。在五个或更多说话人的拥挤对话中,这一问题变得更加困难。我们评估了端到端神经说话人日志模型如何统计说话人数量,并发现其在拥挤录音中系统性低估。标准说话人日志错误率(DER)掩盖了这一缺陷,因为它是基于时长的加权指标,几乎不惩罚被丢弃的低活跃度说话人。因此,我们还报告了Jaccard错误率(JER)和显式计数指标。我们提出了一种门控损失,将说话人存在性与帧活动性耦合。该损失可通过两种方式计算:时长加权或说话人加权,其中说话人加权变体作为正则化器,可减少低估。在真实世界的拥挤录音中,我们的方法明显改善了计数指标,使JER相对降低约6%,DER相对降低约9%,同时不影响稀疏录音的性能。
英文摘要
Speaker diarization must solve two problems: counting how many speakers are present in a given conversation, and assigning speech to each one. This becomes harder in crowded conversations with five or more speakers. We evaluate how end-to-end neural diarization models count speakers, and find systematic under-counting in crowded recordings. The standard diarization error rate (DER) hides this failure, because it is duration-weighted and barely penalizes the dropped, low-activity speakers. We therefore also report the Jaccard error rate (JER) and explicit counting metrics. We propose a gated loss that couples speaker existence with frame activity. This loss can be computed in two ways, duration-weighted or speaker-weighted, and the speaker-weighted variant, used as a regularizer, reduces the under-count. On real-world crowded recordings our method clearly improves the counting metrics, lowering JER by about 6% relative and DER by about 9% relative, while leaving sparse recordings unharmed.
发表机构
- Fano
机构由 AI 辅助整理,请以论文原文为准。