上下文跨越:全双工语音模型与外部LLM后端之间的通信框架
Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends
- Mindlogic(Mindlogic公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出上下文跨越框架,通过实时分块预填充将外部LLM检索信息直接注入全双工语音模型,避免压缩损失,提升全双工基准和问答任务性能。
AI中文摘要:
全双工语音对话模型可以像人类对话的实时动态一样同时进行听和说。对于自然对话而言,实时搜索外部信息的能力也是一项重要能力。许多模型仍局限于参数化知识,无法访问实时信息和执行工具。此外,即使大型语言模型(LLM)检索到信息,许多双工语音模型也是在压缩的潜在空间而非原始文本形式中处理这些信息,这可能导致压缩造成的信息丢失。为解决这一问题,我们提出了上下文跨越(Context Spanning),一种通过实时分块预填充在全双工语音模型与外部LLM后端之间进行信息注入的框架。注入的帧在实时帧预算内的单次前向传播中编码。它将检索到的信息原样输入语音模型,使其能够独立地对信息进行推理并生成响应。通过这种方法,我们的模型在全双工基准测试中取得了高性能,并在问答任务上获得了强劲结果,展示了其对话潜力。上下文跨越表明,外部信息可以直接注入双工语音模型,为双工系统引入了一种简单而强大的新机制。
英文摘要:
Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving them unable to access real-time information and tool execution. Furthermore, even when Large Language Models (LLM) retrieve information, many duplex speech models process it within a compressed latent space rather than in its raw text form, which can lead to information loss from compression. To address this issue, we propose Context Spanning, a framework for information injection between a full-duplex speech model and an external LLM backend via real-time chunked prefill. The injected frame is encoded in a single forward pass inside the real-time frame budget. It feeds the retrieved information to the speech model as-is, enabling it to reason over the information independently and generate responses. With this approach, our model achieves high performance on Full-Duplex benchmarks and strong results on Question Answering tasks, demonstrating its conversation potential. Context Spanning shows that external information can be injected directly into a duplex speech model, introducing a new simple and powerful mechanism for duplex systems.