发表机构
HiThink Research(HiThink研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出HiThink Turn,一种意图感知的流式轮次状态预测器,通过分离回应意图与语义完整性并利用最小意图充分前缀监督,实现了低延迟、高准确的全双工对话轮次控制,显著提升了打断成功率并降低了停止延迟。
AI 中文摘要
全双工对话需要及时且有选择性地处理打断,仅靠话轮结束预测无法实现这一点:完整的语句可能无需回应,而未完成的请求则可能需要打断。为应对这一挑战,我们提出了HiThink Turn,一种意图感知的流式轮次状态预测器,它将回应意图与语义完整性分离,并根据系统播放状态来决策。一个关键贡献是最小意图充分前缀监督,通过LLM判断和语音对齐构建,而在块边界处截断的音频上进行训练则提高了对部分语音的鲁棒性。这些组件支持240毫秒音频块的流式推理,实现了低延迟、准确的全双工轮次控制。实验表明,HiThink Turn在Easy Turn宏平均准确率、Full-Duplex-Bench平均交互率得分(0.933)和非目标语音平均播放恢复率(0.735)上领先于对比方法。此外,意图前缀触发将打断成功率从89%提高到98%,并将平均停止延迟降低了60.9%。
英文摘要
Full-duplex dialogue requires timely yet selective interruption handling, which end-of-turn prediction alone cannot achieve: complete utterances may need no response, while unfinished requests may warrant interruption. To address this challenge, we propose HiThink Turn, an intent-aware streaming turn-state predictor that separates response intent from semantic completeness and conditions decisions on system playback state. A key contribution is minimal intent-sufficient prefix supervision, constructed through LLM judgments and speech alignment, while training on audio truncated at chunk boundaries improves robustness to partial speech. These components support streaming inference with 240-ms audio chunks, enabling low-latency, accurate full-duplex turn control. Experiments show that HiThink Turn leads the compared methods in Easy Turn macro accuracy, Full-Duplex-Bench average interaction rate score (0.933), and non-target-speech average playback resume rate (0.735). Additionally, intent-prefix triggering raises interruption success from 89\% to 98\% and reduces mean stop latency by 60.9\%.