发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出统一框架,将流式多说话人语音识别划分为四种架构策略,基于共享开源模型构建四个系统,评估多说话人准确性、单说话人性能下降、内存和训练复杂度,为不同部署场景提供选型指导。
AI 中文摘要
流式多说话人语音识别是一项具有挑战性的任务,它必须在处理重叠语音、以在线方式对长时间对话进行连贯的长上下文建模的同时,平衡准确性、延迟和效率。我们提出了一个统一框架,根据说话人日志(diarization)与语音识别(ASR)的集成方式,将流式多说话人语音识别划分为四种架构策略。利用一对共享的开源流式语音识别和说话人日志模型作为共同基础,我们构建了四个多说话人语音识别系统,这些系统的区别在于是否采用多个模型实例、微调,或两者兼用。我们在多说话人准确性、单说话人准确性下降程度、内存占用和训练复杂度等方面对这些系统进行了评估。通过这一系统性的架构分析,我们阐明了流式多说话人语音识别的设计空间,并为在不同部署约束下选择最合适的方法提供了实用指导。
英文摘要
Streaming multi-speaker ASR is a challenging task that must balance accuracy, latency, and efficiency while handling overlapping speech and maintaining coherent long-context modeling over extended conversations in an online fashion. We present a unified framework that categorizes streaming multi-speaker ASR into four architectural strategies based on how diarization and ASR are integrated. Using a shared pair of open-source streaming ASR and diarization models as a common foundation, we derive four multi-speaker ASR systems that differ in whether they employ multiple model instances, fine-tuning, or both. We evaluate these systems across multi-speaker accuracy, single-speaker accuracy degradation, memory footprint, and training complexity. Through this systematic architectural analysis, we clarify the design space for streaming multi-speaker ASR and provide practical guidance for selecting the most suitable approach under diverse deployment constraints.
CommentsThis article is accepted to Interspeech 2026