arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02941eess.AScs.LGcs.SD

FASTDIAR:面向流式说话人日志的帧级说话人编码器

FASTDIAR: Frame-level speaker encoder for Streaming Diarization

发表机构KTH皇家理工学院 · 微软
查看机构详情
  • KTH Royal Institute of Technology(KTH皇家理工学院)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

Nikita Torgashov, Okan Köpüklü

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出FASTDIAR,一种帧级因果说话人编码器与在线聚类结合的流式说话人日志系统,通过蒸馏训练,在低重叠基准上实现亚秒延迟的最优准确率,且单CPU线程运行速度比实时快五倍。

中文摘要 AI 辅助

实时对话智能体需要能够在CPU上流式运行且实时处理的说话人日志系统。大多数现有系统采用话语级说话人编码器处理短且高度重叠的音频块,这既浪费计算资源,又使模型针对错误的任务进行优化。我们转而将最先进的说话人识别架构改造为因果帧级编码器,该编码器一次性读取音频流,并每80毫秒从过去两秒的有界音频窗口中输出一个嵌入向量,同时与在线聚类算法配对,该算法根据流自身的相似性对每次更新进行门控。该系统仅通过从话语级教师模型在模拟及域外混合语音上进行蒸馏训练,并使用一组固定的超参数进行评估,在亚秒级延迟下,成为低重叠基准上最准确的流式说话人日志系统;随着说话人数量的增加,其性能下降幅度远小于基于缓存的方法,并且在单CPU线程上的运行速度比实时快五倍。

英文摘要

Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most accurate streaming diarizer on low-overlap benchmarks at sub-second latency, degrades far less than cache-based systems as the number of speakers grows, and runs five times faster than real time on a single CPU thread.

补充信息

↑