arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25195eess.AScs.MA

Qwen-Audio-Agent 技术报告

Qwen-Audio-Agent Technical Report

  • Alibaba Token Foundry, Alibaba Group(阿里巴巴通义实验室,阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Chong Deng, Yunjie Ji, Yuxiang Kong, Xiangang Li, Xu Li, Binbin Zhang, Haina Zhu, Jianheng Zhuo

中文总结 AI 辅助

本文提出Qwen-Audio-Agent,采用前台-后台架构实现全双工语音交互与异步任务执行,通过混合执行在座舱基准上达到91.04%成功率并显著降低延迟。

中文摘要 AI 辅助

我们提出了 Qwen-Audio-Agent,这是一个通过前台-后台架构将全双工语音交互与异步任务执行相结合的框架。前台代理管理对话,并在直接使用工具和委派之间进行选择,而后台代理在单独的上下文中执行委派的任务。编排运行时维护任务状态,协调用户输入和授权的请求,并安排结果返回对话。运行时将语音中断与任务取消分开,并将执行完成与结果交付分开,从而允许在委派工作进行时继续对话。环境事件和持久记忆在会话内和跨会话提供上下文。独立的适配器支持与不同的前台模型、后台代理和客户端集成。我们在桌面助手、智能座舱和语音客服中实例化了该架构。在包含134个案例的内部座舱基准测试中,混合执行的任务成功率达到91.04%,而直接配置和全委派配置分别为72.39%和80.60%。在另一项针对匹配成功轮次的延迟评估中,混合执行相对于这些基线分别将平均任务执行延迟降低了26.73%和30.91%。这些结果支持将直接工具调用用于即时操作和后台委派用于多步骤任务的互补使用。

英文摘要

We present Qwen-Audio-Agent, a harness that combines full-duplex voice interaction with asynchronous task execution through a foreground-background architecture. A Frontend Agent manages dialogue and selects between direct tool use and delegation, while a Backend Agent carries out delegated tasks in a separate context. An Orchestration Runtime maintains task state, coordinates requests for user input and authorization, and schedules the return of results to the conversation. The runtime separates speech interruption from task cancellation and execution completion from result delivery, allowing conversation to continue while delegated work proceeds. Environmental events and persistent memory provide context within and across sessions. Independent adapters support integration with different frontend models, backend agents, and clients. We instantiate the architecture in desktop assistance, intelligent cockpits, and voice customer service. On an in-house cockpit benchmark of 134 cases, mixed execution achieves a task success rate of 91.04%, compared with 72.39% and 80.60% for the direct and all delegated configurations, respectively. In a separate latency evaluation on matched successful turns, mixed execution reduces mean task execution latency by 26.73% and 30.91% relative to these baselines, respectively. These results support the complementary use of direct tool calls for immediate operations and backend delegation for multi-step tasks.

↑