发表机构
NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出前后端架构,通过双工语音转文本前端委派令牌并利用后端LLM执行工具调用,在保持低延迟交互的同时实现高召回率与强代理能力。
AI 中文摘要
全双工语音到语音(S2S)模型提供了自然、低延迟的对话交互,并且若能使用外部工具并完成语音代理任务将受益匪浅。我们提出了一种前后端架构,其中双工语音到文本前端学习发出委派令牌,并将流式ASR转录文本转发给基于文本的后端大语言模型(LLM)以进行工具调用。来自后端的工具调用结果通过轻量级的预填充和重复机制注入回前端,然后使用流式TTS合成给用户。我们的方法在很大程度上保留了常规的双工轮流发言、中断处理和低延迟交互,因为它对前端模型只需进行最小程度的修改。在单轮工具调用评估中,我们的系统实现了92-97%的工具调用召回率、具有竞争力的工具调用预测性能,以及在拒绝不相关调用方面81.2%的准确率。当配备更大的后端(例如Qwen3-235B-A22B)时,我们的系统在Full-Duplex-Bench-V3上取得了与开源和闭源模型相比具有竞争力的结果,并在EVA-Bench上显著优于GPT-realtime-mini和Qwen3-Omni-30B-A3B-Instruct。这些结果表明,后端委派是将自然双工语音交互与强大的代理工具调用能力相结合的一种有效且模块化的方法。
英文摘要
Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent tasks. We propose a frontend-backend architecture where a duplex speech-to-text frontend learns to emit a delegation token and forwards streaming ASR transcripts to a text-based backend LLM for tool calls. Tool-call results from the backend are injected back into the frontend through a lightweight prefill-and-repeat mechanism and then synthesized using streaming TTS to the user. Our approach largely preserves regular duplex turn-taking, interruption handling, and low-latency interaction as it requires minimal modifications to the frontend model. In a single-turn tool-call evaluation, our system achieves 92-97% tool-call recall, competitive tool-call prediction performance, and 81.2% accuracy in rejecting irrelevant calls. When equipped with a larger backend (e.g., Qwen3-235B-A22B), our system achieves competitive results on Full-Duplex-Bench-V3 compared to open and closed source models, and significantly outperforms GPT-realtime-mini and Qwen3-Omni-30B-A3B-Instruct on EVA-Bench. These results demonstrate that backend delegation is an effective and modular approach for combining natural duplex speech interaction with strong agentic tool-call capabilities.
CommentsTo be submitted to ICASSP'27