发表机构
Tsinghua University; ByteDance(清华大学; 字节跳动)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SALMONN-duo,一种受双过程理论启发的自适应双系统全双工语音代理,通过快速系统1与慢速系统2的协调委派,在实时交互中提升知识密集型与多跳推理准确性,并以成本感知强化学习优化性能与后端使用权衡。
AI 中文摘要
全双工语音大语言模型(LLMs)能够实现低延迟、自然的语音交互。然而,现实世界中的代理还必须使用工具并执行深思熟虑的推理操作,这些操作的延迟和计算成本可变,与实时对话的严格时序要求相冲突。为了调和这些需求,我们提出了SALMONN-duo,一种受认知双过程理论启发的自适应双系统语音代理。SALMONN-duo通过将始终开启、快速思考的全双工语音LLM(系统1)与强大的异步慢速思考LLM代理(系统2)配对,将实时交互与深思熟虑的计算分离。除了处理实时交互外,系统1还学习何时直接回答以及何时委派,在后端执行期间保持响应性,并无缝地将返回的信息整合到持续对话中,而不暴露工具痕迹或丢失对话上下文。在单轮口语问答(QA)和多轮对话上的评估表明,自适应委派显著提高了知识密集型和多跳推理问题的准确性,而知识边界感知训练避免了不必要的系统2调用。在定制的τ-Voice版本上,SALMONN-duo进一步展示了其通过多轮交互在现实业务场景中完成环境基础、策略约束任务的能力。最后,成本感知强化学习进一步增强了QA和对话任务中任务性能与后端使用之间的权衡,同时在τ-Voice上以可接受的委派率增加提高了任务成功率和响应安全性。
英文摘要
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conversation. To reconcile these demands, we propose SALMONN-duo, an adaptive dual-system voice agent inspired by dual-process theories of cognition. SALMONN-duo separates real-time interaction from deliberative computation by pairing an always-on, fast-thinking full-duplex speech LLM (system 1) with a powerful asynchronous slow-thinking LLM agent (system 2). Beyond handling real-time interaction, system 1 learns when to answer directly and when to delegate, remaining responsive during backend execution and seamlessly integrating returned information into the ongoing dialogue without exposing tool traces or losing conversational context. Evaluations on single-turn spoken question answering (QA) and multi-turn conversations demonstrate that adaptive delegation substantially improves accuracy on knowledge-intensive and multi-hop reasoning questions, while knowledge-boundary-aware training avoids unnecessary system 2 invocations. On a customized version of $τ$-Voice, SALMONN-duo further demonstrates its ability to complete environment-grounded, policy-constrained tasks through multi-turn interactions in realistic business scenarios. Finally, cost-aware reinforcement learning further enhances the trade-off between task performance and backend usage across the QA and conversation tasks, while improving task success and response safety on $τ$-Voice with an acceptable increase in the delegation rate.