发表机构
Arizona State University(亚利桑那州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对助听器低延迟需求,提出含流式说话人分割与验证的两步法管道,经多工具基准测试与参数优化,17个评估片段中中位数准确率超0.90,为低延迟选择性放大提供概念验证。
AI 中文摘要
受助听器应用的驱动(此类应用中低至10毫秒的延迟都可能产生感知影响),我们提出了一种基于开源预训练模型的实时流式目标说话人识别管道。我们构建了两步法:首先通过低延迟流式说话人 diarization( diarization 指说话人分割)对输入音频进行说话人分割,再针对已注册的目标说话人进行说话人验证。为模拟会话语音同时最小化重叠,我们使用了 This American Life Podcast Transcripts 数据集,并选择主持人作为一致的目标说话人。我们使用 diarization error rate(DER, diarization error rate 指 diarization 错误率)对离线 diarization 进行基准测试,分别采用 Pyannote 和 LIUM 工具,基于基线性能和流式兼容性选择了 Pyannote。随后我们使用 Pyannote 和 TitaNet-Large 进行说话人验证,并生成 ROC 曲线以选择操作区间。我们整合了 Diart 并调整了聚类参数,以降低 DER 同时保持实时运行。我们将 Diart 与 Pyannote 验证模块配对,通过将预测和真实的语音区域转换为100毫秒的二进制掩码来评估系统级性能。在17个评估片段中,该系统在余弦距离阈值为0.7至0.75时,实现了超过0.90的中位数准确率和0.95至0.98的高特异性,为下游低延迟选择性放大提供了实用的概念验证。
英文摘要
We present a real-time pipeline of open source, pretrained models for streaming identification of a target speaker, motivated by hearing-aid applications where latency as low as 10 ms can be perceptible. We formulate a two-step approach in which incoming audio is first segmented by speaker using low-latency streaming diarization, followed by speaker verification against a registered target speaker. To emulate conversational speech while minimizing overlap, we use the This American Life Podcast Transcripts dataset and select the host as a consistent target speaker. We benchmark offline diarization with Pyannote and LIUM using diarization error rate (DER) and select Pyannote based on baseline performance and compatibility with streaming. We then evaluate speaker verification using Pyannote and TitaNet-Large and generate ROC curves to select an operating region. We integrate Diart and tune clustering parameters to reduce DER while maintaining real-time operation. We pair Diart with Pyannote verification and evaluate system-level performance by converting predicted and ground-truth speech regions into 100 ms binary masks. Across 17 evaluation episodes, the system achieves greater than 0.90 median accuracy with high specificity (0.95-0.98) at cosine distance thresholds of 0.7-0.75, demonstrating a practical proof of concept for downstream low-latency selective amplification.
Comments8 pages, 4 figures