arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiffuPlex:通过滚动掩码扩散加速全双工口语对话模型

DiffuPlex: Accelerating Full-Duplex Spoken Dialog Models via Rolling Masked Diffusion

Heeseung Kim

arXiv 2610.12214首次发表:更新:

发表机构

University of Seoul(首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DiffuPlex是一种滚动掩码扩散框架,通过预测多帧减少全双工口语对话模型的顺序计算,两种策略均实现显著速度提升,且基本保留交互与语言能力。

AI 中文摘要

现有的全双工口语对话模型支持同时听和说,但细粒度模型仍在每个交互帧以自回归方式推进其主干网络。我们提出DiffuPlex,一种滚动掩码扩散框架,通过单次主干网络前向传播预测多个未来用户和助手帧,减少这种顺序计算。DiffuPlex仅使用每个预测未来的可信前缀,同时以原始帧速率继续交互。当用户语音到达时,它会检查对应的用户预测,若交互出现偏差,则保留已播放的助手内容,仅修改未播放的未来部分。我们在同一预测器上考虑两种推理策略:DiffuPlex-LISTEN在预测助手静音时使用多个未来帧,而DiffuPlex-SPEAK还可使用预测的助手语音。在全双工交互和口语语言评估中,DiffuPlex大幅减少了顺序主干计算,同时基本保留了交互行为和通用能力。DiffuPlex-LISTEN和DiffuPlex-SPEAK分别实现了1.46倍和1.59倍的部署路径挂钟速度提升,以及1.61倍和1.80倍的Core LM速度提升,所有测量的主干网络调用均在80ms交互间隔内完成。人工评估显示,LISTEN策略保留了语音自然度和对话质量,而SPEAK策略保留了对话质量,但语音自然度略有下降。

英文摘要

Recent full-duplex spoken dialog models enable simultaneous listening and speaking, but fine-grained models still advance their backbone autoregressively at every interaction frame. We introduce DiffuPlex, a rolling masked diffusion framework that reduces this sequential computation by predicting multiple future user and assistant frames in a single backbone wake. DiffuPlex consumes only a confident prefix of each predicted future while interaction continues at the original frame rate. As user speech arrives, it checks the corresponding user predictions and, when the interaction diverges, preserves already played assistant content while revising only the unplayed future. We consider two inference policies over the same predictor: DiffuPlex-LISTEN consumes multiple future frames when they predict assistant silence, whereas DiffuPlex-SPEAK can also consume predicted assistant speech. Across full-duplex interaction and spoken-language evaluations, DiffuPlex substantially reduces sequential backbone computation while largely preserving interaction behavior and general capability. DiffuPlex-LISTEN and DiffuPlex-SPEAK achieve $1.46\times$ and $1.59\times$ deployment-path wall-clock speedups and $1.61\times$ and $1.80\times$ Core LM speedups, with all measured backbone invocations completing within the 80ms interaction interval. Human evaluation shows that LISTEN preserves speech naturalness and conversational quality, while SPEAK retains conversational quality with some degradation in speech naturalness.

Comments41 pages, 11 figures, 18 tables. Preprint, under review. Project page: https://diffuplex.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑