arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁说了什么,它会被记住吗?评估跨会议中的持久说话人归属

Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings

Shantanu Vispute, Aditya Mishra, Siddhartha Saxena

arXiv 2609.39344首次发表:更新:

发表机构

Foyer(Foyer)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SI-cpWER指标评估跨会议持久说话人归属,对比商业与学术系统,ThyVoice在CHiME-8上表现最优,SI-cpWER为47.13,优于商业系统,强调持久归属需直接评估。

AI 中文摘要

用作长期记忆的语音转录必须同时保留词语和稳定的说话人身份。现有的会议转录指标要么忽略说话人,要么在每个录音中独立地重新映射匿名说话人,因此它们无法衡量同一个人是否在跨会议中保持同一身份。我们使用说话人识别cpWER(SI-cpWER)来评估持久说话人归属,该指标在全局说话人ID分配下对整个语料库进行评分。该基准涵盖了五个商业化的“先分割再识别”级联系统、两个开放学术基线以及ThyVoice,在完整的129场会议CHiME-8 NOTSOFAR评估集上以干净和噪声增强形式进行测试,同时还包括CHiME-6。ThyVoice是我们的端到端参考系统;它修复重叠并门控用于创建和更新声纹的证据。要求持久身份改变了商业排名:在所有三种条件下,ThyVoice记录的SI-cpWER低于每个评估的商业级联系统,并且在完整面板中具有最低平均值,为47.13,而下一个系统为54.75。补充的词汇、分割、每录音归属和说话人聚类诊断表征了最终归属记录中的上游错误表面。这些结果说明了为什么在跨时间重用对话的系统中必须直接评估持久归属。

英文摘要

Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently in each recording, so they cannot measure whether the same person retains one identity across meetings. We evaluate persistent speaker attribution with Speaker Identified cpWER (SI-cpWER), which scores a corpus under one global speaker-ID assignment. The benchmark covers five commercial diarize-then-identify cascades, two open academic baselines, and ThyVoice on the full 129-meeting CHiME-8 NOTSOFAR evaluation set in clean and noiseaugmented form, plus CHiME-6. ThyVoice is our end-to-end reference system; it repairs overlap and gates the evidence used to create and update voiceprints. Requiring persistent identity changes the commercial ranking: ThyVoice records lower SI-cpWER than every evaluated commercial cascade in all three conditions and the lowest mean in the full panel, 47.13 versus 54.75 for the next system. Complementary lexical, diarization, per-recording attribution, and speaker-clustering diagnostics characterize upstream error surfaces in the final attributed record. These results show why persistent attribution must be evaluated directly in systems that reuse conversations across time.

CommentsA short version is accepted at IEEE SLT 2026, Demo Track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑