SteerablePlex:我们能否操控全双工模型?
SteerablePlex: Can We Steer Full-Duplex Models?
AI总结:
针对全双工语音模型难以控制的问题,提出SimIF-Bench基准,基于GDPO训练方案构建SteerablePlex,结合异步后端语言模型得到更可控的全双工用户模拟器,性能优于现有开源模型和GPT-Realtime。
AI中文摘要:
全双工语音模型可同时听和说,支持自然交互,但随着对话历史增长,其控制难度不断提升。当将其用作用户模拟器时,这种控制缺失会导致它们偏离规定场景,产生不可靠的评估结果。我们推出SimIF-Bench(模拟器指令遵循基准),用于评估对话模型是否保持在规定场景内并按要求顺序完成多个目标。该基准显示,当前开源全双工模型难以遵循此类约束。随后,我们推出基于Group Reward-Decoupled Normalization Policy Optimization(GDPO)的训练方案,使全双工模型在持续对话中遵循文本指令,同时保持其轮次转换能力。通过将生成的SteerablePlex连接至异步后端语言模型,该模型可监控对话并在需要时提供指令,我们构建了一个更可控的全双工用户模拟器,其遵循多阶段约束的可靠性优于现有开源模型和GPT-Realtime。
英文摘要:
Full-duplex speech models can listen and speak simultaneously, enabling natural interaction, but become increasingly difficult to control as the conversation history grows. When used as user simulators, this lack of control can cause them to deviate from prescribed scenarios and produce unreliable evaluation outcomes. We introduce SimIF-Bench (Simulator Instruction-Following Benchmark), which evaluates whether a conversational model stays within a prescribed scenario and completes multiple goals in the required order. The benchmark reveals that current open-source full-duplex models struggle to follow such constraints. We then introduce a Group Reward-Decoupled Normalization Policy Optimization (GDPO)-based training recipe that enables a full-duplex model to follow textual instructions during an ongoing conversation while maintaining its turn-taking ability. By connecting the resulting SteerablePlex to an asynchronous backend language model that monitors the conversation and provides instructions when needed, we build a more controllable full-duplex user simulator that follows multi-stage constraints more reliably than existing open-source models and GPT-Realtime.