发表机构
Tsinghua University; Huawei Technologies Co., Ltd.(清华大学; 华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对全双工对话系统的现有模型缺陷,提出基于LLM的TurnFSM框架,通过内化状态机逻辑统一流式语义VAD与语句级拒绝,性能优于基线且具竞争力。
AI 中文摘要
全双工语音助手需在说话时持续监听,在低延迟、资源受限的流式场景下处理用户中断。现有端到端全双工模型在语音域适配后会损害与推理相关的能力,而级联流水线会引入额外推理开销和手工控制逻辑。我们提出TurnFSM,一种基于LLM的状态预测框架,将轮次控制内化为显式有限状态转换,统一流式语义语音活动检测(VAD)与语句级拒绝。TurnFSM将提交与拒绝分解为串行决策过程,减少多任务干扰,同时保持与单任务模型相当的性能。我们进一步引入一阶状态转换机制,在训练期间强制仅依赖前一状态,通过标准因果掩码和原始LLM位置编码实现紧凑推理,避免历史状态令牌积累和不必要的逐步状态生成。实验结果显示,TurnFSM始终优于二元头基线,且与任务特定模型相比仍具竞争力。
英文摘要
Full-duplex voice assistants must continuously listen while speaking, handling user interruptions under low-latency and resource-constrained streaming conditions. Existing end-to-end full-duplex models can compromise reasoning-related capabilities after speech-domain adaptation, whereas cascaded pipelines introduce extra inference overhead and handcrafted control logic. We propose TurnFSM, an LLM-based state prediction framework that internalizes turn control as explicit finite-state transitions, unifying streaming semantic VAD and utterance-level rejection. TurnFSM decomposes submission and rejection into a serial decision process, reducing multi-task interference while maintaining performance comparable to single-task models. We further introduce a first-order state transition mechanism that enforces the dependency on only the previous state during training, enabling compact inference with the standard causal mask and original LLM positional encoding while avoiding historical state-token accumulation and unnecessary step-by-step state generation. Experimental results show that TurnFSM consistently outperforms the binary-head baseline and remains competitive with task-specific models.