arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Qwen-Audio-3.1-Realtime:迈向可靠的智能体语音交互

Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

Lujia Bao, Qian Chen, Luyao Cheng, Chong Deng, Yuxiang Kong, Xiangang Li, Xu Li, Jiaqing Liu, Chao-Hong Tan, Haoyu Wang, Wen Wang, Xilou Wang, Haoxiang Xu, Junhao Xu, Liang Yi, Binbin Zhang, Qinglin Zhang, Qiquan Zhang

arXiv 2609.25176首次发表:更新:

AI 中文总结

Qwen-Audio-3.1-Realtime通过思考、行动、说话与协调机制,结合M²-OPD蒸馏和GRPO优化,将任务成功率提升至82.0%,并大幅降低背景语音干扰响应率。

AI 中文摘要

实时语音助手必须能够对不断变化的请求进行推理、执行操作并遵循对话规则。Qwen-Audio-3.1-Realtime通过“思考、行动、说话与协调”将这些要求整合在一起。“思考”将Core-Cocktail监督微调与多模态和多教师在线策略蒸馏(M$^{2}$-OPD)相结合,以迁移语言能力并开发原生音频技能。“行动”利用自进化的可执行环境和多粒度回放进行组相对策略优化(GRPO),教会模型使用工具、解释反馈并完成任务。“说话与协调”则对齐助手说话或行动的时机、方式及是否执行。我们评估了音频推理、多语言理解、工具使用、对话行为、全双工交互和安全性。与Qwen-Audio-3.0-Realtime相比,3.1在我们对$\ au$-Voice的半双工语音到文本改编中,将整体任务成功率从78.4%提升至82.0%。在语音到语音的全双工基准Full-Duplex-Bench v1.5上,对背景语音的响应率从73.0%降至13.0%。我们还展示了一个独立的Voice Harness原型,以前景使用Qwen-Audio-3.0-Realtime,通过前景-背景协调和记忆将口语交互扩展到持久任务。

英文摘要

Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these requirements together through Think, Act, and Speak and Coordinate. Think combines Core-Cocktail supervised fine-tuning with Multimodality and Multi-Teacher On-Policy Distillation (M$^{2}$-OPD) to transfer language capabilities and develop native audio skills. Act uses self-evolving executable environments and multi-granularity rollouts for Group Relative Policy Optimization (GRPO), teaching the model to use tools, interpret feedback, and complete tasks. Speak and Coordinate aligns whether, when, and how the assistant speaks or acts. We evaluate audio reasoning, multilingual understanding, tool use, conversational behavior, full-duplex interaction, and safety. Compared with Qwen-Audio-3.0-Realtime, 3.1 raises overall task success from 78.4% to 82.0% on our half-duplex speech-to-text adaptation of $τ$-Voice. On speech-to-speech Full-Duplex-Bench v1.5, the response rate to background speech falls from 73.0% to 13.0%. We also present a separate Voice Harness prototype, using Qwen-Audio-3.0-Realtime as its foreground, that extends spoken interaction to persistent tasks through foreground--background coordination and memory.

Comments25 pages, technical report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑