arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaptDuplex:从静态到自适应的全双工口语对话

AdaptDuplex: from static to adaptive full-duplex spoken dialogue

Zhiyang Zhou, Yingxin Shang, Zhou Wang, Hongwei Cai, Weixu Wang, Shuran Zhou, Shuofeng Zhao, Wenke Fan, Qingxiang Guo, Dawei Yang, Lin Yang, Yang Song

arXiv 2609.29217首次发表:更新:

发表机构

Zuoyebang Education Technology(作业帮教育科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对全双工口语对话缺乏自适应决策机制的问题,提出AdaptDuplex,通过令牌级协议、动态窗口预测及渐进式训练升级Qwen3-Omni,在多个基准上超越现有模型,取得最高得分72.9。

AI 中文摘要

全双工口语对话要求在亚秒级延迟下同时进行听和说,并需应对时刻变化的对话时机与认知需求。然而,当前模型大多采用静态运行点,缺乏自适应决策的系统性机制。我们提出AdaptDuplex,通过三层协同设计为Qwen3-Omni升级了该机制。一个紧凑的令牌级协议将每个窗口表示为规范序列,通过有界文本领先于语音来训练双流对齐,并将每个行为决策暴露为显式令牌,以便通过logits偏置实现免训练的运行时控制。自适应机制动态预测离散窗口时长,并按需通过非阻塞认知整合和多路外部推理增强直接响应。一个渐进式流水线通过三阶段Thinker课程引入这些行为,随后进行仅Talker和联合SFT,并以GRPO作为进一步增量。在Full-Duplex-Bench v1和v1.5上,AdaptDuplex在大多数可比较的轮流发言、重叠行为和时序指标上优于DuplexOmni和MiniCPM-o 4.5,在交互决策和响应时序方面均获得提升。在人工录制的HumDial-FDBench上,它取得了所比较双工模型中最高的最终得分(72.9)。

英文摘要

Full-duplex spoken dialogue requires simultaneous listening and speaking at sub-second latency, under conversational timing and cognitive demands that change moment to moment. Yet current models mostly impose static operating points, lacking a systematic mechanism for adaptive decisions. We present AdaptDuplex, which extends Qwen3-Omni with such a mechanism, co-designed across three layers. A compact token-level protocol represents every window as a canonical sequence, trains dual-stream alignment through a bounded text lead over speech, and exposes every behavioral decision as an explicit token for training-free runtime control via logits bias. Adaptive mechanisms dynamically predict among discrete window durations and augment direct response as needed with non-blocking cognitive consolidation and multi-flight external reasoning. A progressive pipeline introduces these behaviors through a three-stage Thinker curriculum, then Talker-only and joint SFT, with GRPO as a preliminary increment. On Full-Duplex-Bench v1 and v1.5, AdaptDuplex outperforms DuplexOmni on 18 of 21 comparable turn-taking, overlap-behavior, and timing metrics and MiniCPM-o 4.5 on 19 of 22, with gains in both interaction decisions and response timing. On human-recorded HumDial-FDBench, it attains the top Final score (69.6) of the compared duplex models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑