arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过广义风格感知全双工框架实现主动口语话轮

Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework

Tianrui Pan, Qinglin Zhang, Chong Deng, Luyao Cheng, Qian Chen, Wen Wang, Jie Tang, Gangshan Wu, Jie Liu

arXiv 2608.28630首次发表:更新:

发表机构

Token Foundry, Alibaba Group(阿里巴巴集团Token Foundry)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究构建含LPS-TC控制器、WildTurn数据集及两级评估方案的风格感知全双工框架,结合半双工/全双工模型后,可实现更自然类人的主动口语交互,时机与响应质量表现优异。

AI 中文摘要

与半双工对话系统(系统需等待用户话轮完成后再响应)不同,自然全双工对话系统要求智能体实时主动行动,包括及时打断和反馈通道,这带来了一个关键挑战:在不牺牲响应质量的前提下优化话轮时机。为解决现实主动话轮转换的局限性,我们构建了一个包含三个关键组件的广义风格感知全双工框架。首先,我们提出LPS-TC(轻量型主动口语话轮控制器),支持即插即用集成,其具备覆盖反应式和主动式话轮行为的细粒度动作空间,可使半双工模型获得全双工能力,并为现有全双工模型提供更优的时机控制。其次,我们构建WildTurn,这是一个大规模真实英语数据集,包含约2981小时经筛选的多轮立体声对话,来自面对面和电话对话,标注有5种话轮转换风格和5种反馈通道风格;在WildTurn上训练的LPS-TC展现出现有静态全双工基准未捕捉到的丰富口语动态。第三,我们引入两级评估方案,在现实流约束下评估块级时机精度和话轮级交互质量。我们将LPS-TC与Qwen2.5-Omni等半双工模型、Freeze-Omni等全双工模型结合的实验,展示了其在时机适宜性和响应质量上的优越性能;该框架还展现出细粒度风格可控性和强泛化性,可实现更自然、更类人的口语交互。

英文摘要

Compared with half-duplex dialogue systems where the system waits for user turn completion before it responds, natural full-duplex dialogue systems require agents to act proactively in real time, including timely interruptions and backchannels. This creates a key challenge: improving turn timing without sacrificing response quality. To address limitations in realistic proactive turn-taking, we build a generalized style-aware full-duplex framework with three key components. Firstly, we propose LPS-TC, a Lightweight Proactive Speech Turn Controller for plug-and-play integration. It features a fine-grained action space covering both reactive and proactive turn behaviors, enabling half-duplex models with full-duplex capabilities and enhancing existing full-duplex models with superior timing control. Secondly, we construct WildTurn, a large-scale, real-world English dataset containing approximately 2,981 hours of filtered multi-turn stereo conversations from face-to-face and telephone conversations, annotated with five turn-taking and five backchanneling styles. Trained on WildTurn, LPS-TC exhibits rich spoken dynamics that are not captured by existing static full-duplex benchmarks. Thirdly, we introduce a two-tier evaluation scheme that assesses both chunk-level timing precision and turn-level interaction quality under realistic streaming constraints. Our experiments, integrating LPS-TC with half-duplex models like Qwen2.5-Omni and full-duplex models like Freeze-Omni, showcase its superior performance in timing appropriateness and response quality. Our framework also demonstrates fine-grained style controllability and strong generalizability, enabling more natural and human-like spoken interactions.

CommentsAccepted by ACM MM 2026

DOI:10.1145/3767308.3835384

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑