FD-VAD:面向流式全双工语音的语义端点检测
FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech
AI总结:
FD-VAD提出一种无需ASR的流式语义端点检测方法,直接基于音频窗口判断继续或停止,在TurnBench上取得最优EOT召回率,实现高效全双工语音交互。
AI中文摘要:
在全双工语音交互中,自然的轮流说话需要从部分语音中判断停顿是表示犹豫还是完整的对话意图。声学语音活动检测缺乏这种语义信息,而基于级联ASR的端点检测则引入了转录依赖和额外的处理阶段。我们将语义端点检测表述为一个因果音频-语言推理任务,并引入FD-VAD,一种无需ASR的流式端点检测器,它将有界的因果音频窗口直接映射为继续/停止决策。FD-VAD结合了冻结的语音编码器、轻量级模态适配器和参数高效适配的语言模型,使用最后块训练目标进行流式推理。我们进一步引入了置信度门控的端点承诺来控制打断与延迟,以及边界聚焦的硬负样本采样,以改善在模糊转折边界附近的决策。在领域内和对话评估中,FD-VAD优于强大的流式和非流式语义转折分类器,并在TurnBench开发集上以零样本设置实现了最高的EOT召回率$0.853$(在FP<=0.10时)。这些结果表明,语义端点检测可以直接从流式音频中执行,无需中间ASR或对话状态跟踪。
英文摘要:
Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages. We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference. We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set $0.853$ (at FP<=0.10) in a zero-shot setting. These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.