arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推动流式多说话人语音识别边界:架构权衡的系统性研究

Pushing the Boundaries of Streaming Multi-Speaker ASR: A Systematic Study of Architectural Trade-offs

Taejin Park, Ivan Medennikov, Kunal Dhawan, Weiqing Wang, Jagadeesh Balam, Boris Ginsburg

arXiv 2609.10265首次发表:更新:

发表机构

NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出统一框架,将流式多说话人语音识别划分为四种架构策略,基于共享开源模型构建四个系统,评估多说话人准确性、单说话人性能下降、内存和训练复杂度,为不同部署场景提供选型指导。

AI 中文摘要

流式多说话人语音识别是一项具有挑战性的任务,它必须在处理重叠语音、以在线方式对长时间对话进行连贯的长上下文建模的同时,平衡准确性、延迟和效率。我们提出了一个统一框架,根据说话人日志(diarization)与语音识别(ASR)的集成方式,将流式多说话人语音识别划分为四种架构策略。利用一对共享的开源流式语音识别和说话人日志模型作为共同基础,我们构建了四个多说话人语音识别系统,这些系统的区别在于是否采用多个模型实例、微调,或两者兼用。我们在多说话人准确性、单说话人准确性下降程度、内存占用和训练复杂度等方面对这些系统进行了评估。通过这一系统性的架构分析,我们阐明了流式多说话人语音识别的设计空间,并为在不同部署约束下选择最合适的方法提供了实用指导。

英文摘要

Streaming multi-speaker ASR is a challenging task that must balance accuracy, latency, and efficiency while handling overlapping speech and maintaining coherent long-context modeling over extended conversations in an online fashion. We present a unified framework that categorizes streaming multi-speaker ASR into four architectural strategies based on how diarization and ASR are integrated. Using a shared pair of open-source streaming ASR and diarization models as a common foundation, we derive four multi-speaker ASR systems that differ in whether they employ multiple model instances, fine-tuning, or both. We evaluate these systems across multi-speaker accuracy, single-speaker accuracy degradation, memory footprint, and training complexity. Through this systematic architectural analysis, we clarify the design space for streaming multi-speaker ASR and provide practical guidance for selecting the most suitable approach under diverse deployment constraints.

CommentsThis article is accepted to Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑