HiPLEX: 全双工语音语言模型的分层策略分解
HiPLEX: Hierarchical Policy Factorization for Full Duplex Speech Language Models
浏览论文内容
中文总结 AI 辅助
HiPLEX通过将全双工语音模型策略分解为时序控制和内容生成两个因子,利用强化学习分别优化,从而提升实时对话中的轮流与反馈协调能力。
中文摘要 AI 辅助
随着人机交互日益对话化,能够进行自然实时对话的全双工语音语言模型变得越来越重要。除了生成合适的回应外,这些模型还必须实时协调轮流说话、反馈信号(backchanneling)和话语权管理。强化学习(RL)提供了一种通过直接反馈交互结果来优化这些行为的方式。然而,现有的RL方法要么将时序反馈应用于令牌策略,要么优化语义内容,使得时序和内容的联合改进问题仍未解决。我们提出了HiPLEX,一个RL框架,它将预训练的全双工文本策略分解为一个决定何时发出内容的控制策略和一个决定发出什么内容的条件内容策略。第一个因子在'pad'、'epad'和'con'之间选择。第二个因子仅在选择'con'时选择一个令牌。这种层次结构描述了每个帧内的条件动作,并使用模型现有的文本头。我们通过从生成的语音片段中派生的事件因果掩码将时序优势路由到令牌组因子,并将LLM评判的语义优势路由到条件内容因子。在Full-Duplex-Bench v1上的三个Moshi种子中,与GRPO相比,HiPLEX在自然用户停顿和反馈机会期间降低了接管率,并缩短了中断后响应延迟,同时保持了相当的评判性中断响应质量。在Moshi和PersonaPlex上,HiPLEX比GRPO更好地匹配了合并的人类轮流时序和反馈率边际分布。
英文摘要
As human--AI interactions become more conversational, full-duplex speech language models capable of natural real-time dialogue are growing in importance. Beyond generating appropriate responses, these models must coordinate turn-taking, backchanneling, and floor management in real time. Reinforcement learning (RL) provides a way to refine these behaviors through direct feedback on interaction outcomes. However, existing RL methods either apply timing feedback to a token policy or optimize semantic content, leaving the joint improvement of timing and content unresolved. We introduce HiPLEX, an RL framework that factorizes a pretrained full-duplex text policy into a control policy that decides when to emit content and a conditional content policy that decides what to emit. The first factor selects among 'pad', 'epad', and 'con'. The second selects a token only when 'con' is chosen. This hierarchy describes conditional actions within each frame and uses the model's existing text head. We route timing advantages to the token-group factor through event-causal masks derived from generated speech episodes, and route an LLM-judge semantic advantage to the conditional content factor. Across three Moshi seeds on Full-Duplex-Bench v1, HiPLEX reduces takeover rates during natural user pauses and backchannel opportunities, and shortens post-interruption response latency relative to GRPO, while maintaining comparable judged interruption-response quality. On Moshi and PersonaPlex, HiPLEX better matches pooled human turn-timing and backchannel-rate marginals than GRPO.
发表机构
- Qualcomm AI Research(高通人工智能研究院)
- KAIST AI(韩国科学技术院人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。