arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向准确的说话人清单的在线说话人 diarization(说话人 diarization 指说话人分割与识别)

Toward accurate speaker inventories for online speaker diarization

Youngki Kwon, Hee-Soo Heo, Minjae Lee, Bong-Jin Lee

arXiv 2610.11278首次发表:更新:

AI 中文总结

该研究针对在线说话人 diarization 中现有系统的说话人数量误差问题,改进基于嵌入的跟踪器的注册触发机制,在9个公开数据集上显著降低宏观说话人数量误差与 DER,在 VoxSRC-23 上同时获最优 DER 和说话人数量精度。

AI 中文摘要

在线说话人 diarization 系统会维护一个不断增长的说话人清单,应用程序可直接读取该清单:实时转录文本会在说话人出现时对其进行枚举。现有系统要么预先固定清单的大小,导致超出该大小的说话人被合并;要么仅基于距离增长清单,导致噪声和话轮边界将同一说话人拆分。在9个公开数据集上按说话人数量评分时,两种失败情况均出现在 diarization error rate(DER, diarization 错误率)未报告问题的地方:在 VoxSRC-23 数据集上,固定容量系统返回的说话人数量少了44%-46%,而具有最低 DER 的无约束跟踪器返回的说话人数量多了63%。我们基于该基于嵌入的跟踪器进行改进,仅修改触发注册的条件:候选说话人池将未注册的嵌入分开,仅当某个候选收集到足够多相互兼容的嵌入时才注册该说话人;提交门控标签分配会在候选未决时暂存输出,确保说话人在确认前不会出现在输出中。在9个数据集上,宏观说话人数量误差从参考值的225%降至21%,宏观 DER 从17.59%降至15.95%,所有输出的平均额外延迟为0.109秒;在 VoxSRC-23 数据集上,该跟踪器在对比系统中同时获得最低 DER 和最准确的说话人数量。

英文摘要

An online speaker diarizer maintains a growing inventory of speakers, and applications read it directly: a live transcript enumerates them as they appear. Existing systems either fix the inventory's size in advance, so speakers beyond that size are merged, or grow it on distance alone, so noise and turn boundaries split speakers. Scored on speaker count over nine public datasets, both failures appear where DER reports neither: on VoxSRC-23 the fixed-capacity systems return 44-46% too few speakers and the unbounded tracker with the lowest DER 63% too many. We build on that embedding-based tracker and change only what triggers a registration. A candidate speaker pool keeps unregistered embeddings apart and registers a speaker only once one candidate has gathered enough mutually compatible embeddings; commit-gated label assignment holds an output while its candidate is undecided, so no speaker appears in the output before it is confirmed. Across the nine datasets the macro speaker-count error falls from 225% to 21% of the reference and macro DER from 17.59% to 15.95%, at a mean added latency of 0.109s over all outputs; on VoxSRC-23 the tracker obtains both the lowest DER and the most accurate speaker count of the systems compared.

Comments5 pages, 1 figure, submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑