发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出在双工语音到语音模型中并行添加轻量级ASR头,实现流式用户转录,在保持全双工对话能力的同时,取得10.21%的流式WER,并开源代码。
AI 中文摘要
全双工语音到语音(S2S)模型通过允许同时进行听和说,实现了自然的对话式人工智能。然而,这些模型通常缺乏固有的用户语音转录能力,而这一能力对于对话记录、无障碍功能和质量监控等应用至关重要。在这项工作中,我们提出了一种高效方法,通过在与智能体文本头并行的位置引入一个轻量级自动语音识别(ASR)头,为现有的双工S2S模型添加流式ASR能力。我们的方法仅需极少的额外参数,且无需对基础S2S模型进行重大架构更改,从而在保持全双工对话能力(包括轮流发言和打断处理)的同时,实现实时的用户转录。实验结果表明,我们的方法在双工S2S框架内,在HuggingFace Open ASR排行榜上实现了10.21%的流式平均词错误率(WER)。此外,我们展示了相同架构作为独立流式ASR模型训练时,与当前最先进(SOTA)模型相比,取得了具有竞争力的结果(WER为7.73%)。我们将开源我们的训练和推理代码,以促进联合流式ASR和S2S建模的进一步研究。
英文摘要
Full-duplex speech-to-speech (S2S) models enable natural conversational AI by allowing simultaneous listening and speaking. However, these models typically lack inherent user speech transcription, which is essential for applications such as conversation logging, accessibility features, and quality monitoring. In this work, we propose an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head. Our approach requires minimal additional parameters and no significant architectural changes to the base S2S model, enabling real-time user transcription while preserving full-duplex conversational capabilities including turn-taking and barge-in handling. Experimental results demonstrate that our method achieves streaming average WER of 10.21% on the HuggingFace Open ASR Leaderboard within the duplex S2S framework. Additionally, we show that the same architecture trained as a standalone streaming ASR model achieves competitive results (7.73% WER) compared to current SOTA models. We will open-source our training and inference code to facilitate further research in joint streaming ASR and S2S modeling.