SteerDuplex:可操控的全双工语音对话模型
SteerDuplex: Steerable Duplex Speech Dialogue Models
浏览论文内容
中文总结 AI 辅助
针对全双工语音对话模型缺乏可操控性的问题,提出SteerDuplex模型和SteerBench基准,通过微调与两阶段强化学习,显著提升音频操控与打断处理能力。
中文摘要 AI 辅助
全双工语音对话模型支持低延迟的轮流说话、打断处理和反馈语,但一个关键能力仍未得到充分探索:可操控性,即根据用户指令,在语气、人格、语速和语音风格等属性上可靠地转变对话行为的能力。我们引入了一个基于文本和音频的可操控性分类体系,该体系揭示了当前全双工模型中的重大不足。为解决这一不足,我们提出了SteerDuplex,一个基于Moshi的全双工语音模型,在自然对话和针对指令遵循、语音表达、推理及双工交互的合成对话上进行微调。我们进一步应用了带有混合奖励的两阶段强化学习(RL),结合可验证的交互检查和基于评判者的语义反馈,以改善时序和响应连续性。为评估全双工语音可操控性,我们引入了SteerBench,一个包含390个语音提示和1,067个人工编写的二元音频和文本评分标准的基准,涵盖语气、人格、风格/口音和速度/长度。在SteerBench上,监督训练将音频操控平均通过率相较于最强评估的开放基线提高了44.5个百分点。在Audio MultiChallenge上,任务平均通过率相较于其最强评估的开放基线提高了7个百分点。RL进一步将源清晰打断响应从72.5%提高到82.5%,并将合成暂停插入从26.5%降低到9%。操控和综合任务得分保持相当或更高,而奖励探测揭示了通过不完整响应进行奖励黑客的行为。我们的模型和基准支持对语音可操控性的系统性研究,奖励分析表明为何时序增益必须与响应完整性一起评估。
英文摘要
Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability to reliably shift conversational behavior along attributes such as tone, persona, speaking rate, and voice style in response to user instructions. We introduce a taxonomy of text- and audio-based steerability that identifies substantial gaps in current full-duplex models. To address this gap, we introduce SteerDuplex, a Moshi-based full-duplex speech model fine-tuned on natural conversations and synthetic dialogues targeting instruction following, vocal delivery, reasoning, and duplex interaction. We further apply two-stage reinforcement learning (RL) with hybrid rewards, combining verifiable interaction checks and judge-based semantic feedback to improve timing and response continuity. To evaluate full-duplex spoken steerability, we introduce SteerBench, a benchmark with 390 spoken prompts and 1,067 human-authored binary audio and text rubrics spanning tone, persona, style/accent, and speed/length. On SteerBench, supervised training improves audio-steering average pass rate by 44.5 percentage points over the strongest evaluated open baseline. On Audio MultiChallenge, task average pass rate improves by 7 points over its strongest evaluated open baseline. RL further raises source-clean interruption response from 72.5% to 82.5% and reduces synthetic pause barge-in from 26.5% to 9%. Steering and aggregate task scores remain comparable or higher, while reward probes reveal reward hacking through incomplete responses. Our model and benchmark support systematic research on spoken steerability, with reward analysis showing why timing gains must be evaluated alongside response completeness.
发表机构
- Scale AI
- University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。