NemotronLabs VoiceChat:具备工具调用能力的开源全双工语音到语音模型
NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities
浏览论文内容
中文总结 AI 辅助
该论文提出开源全双工语音到语音模型NemotronLabs VoiceChat,通过流式架构整合语音编解码、RNN-T转录和工具调用,在多项基准上实现低接管率、高响应质量及82.5%工具选择F1,证明全双工交互与外部工具可统一集成。
中文摘要 AI 辅助
我们推出了NemotronLabs VoiceChat,这是一个具备原生工具调用能力的开源全双工语音到语音模型。NemotronLabs VoiceChat结合了流式语音编码器和仅解码器语言模型,并配有用于智能体文本和结构化函数调用的并行专用输出流、用于增量式用户转录的辅助RNN-T分支,以及流式TTS解码器。这一设计使模型能够在统一的流式架构中实现聆听、转录、推理、调用工具和说话,同时保持自然对话所需的时间行为。在Full-Duplex-Bench 1.0上,NemotronLabs VoiceChat在所评估的开源权重系统中实现了最低的停顿处理接管率,在用户打断后实现100%的接管,并在打断后响应质量上获得4.33/5的评分。在Full-Duplex-Bench 1.5上,它在用户反馈后93%的情况下恢复响应。NemotronLabs VoiceChat在VoiceBench上获得55.1的归一化平均分,并在Full-Duplex-Bench 3.0(FDB 3.0)上实现82.5%的工具选择F1分数,而参数准确性和端到端工具执行仍有改进空间。这些结果表明,全双工交互、语音识别与生成、通用语言能力以及外部工具使用可以集成在单个开源语音到语音模型中,而不会牺牲实时对话行为。
英文摘要
We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming architecture while preserving the temporal behavior required for natural conversation. On Full-Duplex-Bench 1.0, NemotronLabs VoiceChat achieves the lowest pause-handling takeover rates among evaluated open-weight systems, 100\% takeover following user interruptions, and a 4.33/5 post-interruption response-quality score. On Full-Duplex-Bench 1.5, it resumes its response after user backchannels in 93\% of cases. NemotronLabs VoiceChat obtains a 55.1 normalized average on VoiceBench and, on Full-Duplex-Bench 3.0 (FDB 3.0), achieves 82.5\% tool-selection F1, while argument accuracy and end-to-end tool execution remain areas for improvement. These results demonstrate that full-duplex interaction, speech recognition and generation, general language capabilities, and external tool use can be integrated in a single open speech-to-speech model without sacrificing real-time conversational behavior.
发表机构
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。