arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少即是多:哪些语音方面驱动话轮结束检测

Less can be More: What Aspects of Speech Drive End-of-Turn Detection

Rini Sharon, Manickavela A, Kadri Hacioglu, Andreas Stolcke

arXiv 2609.11066首次发表:更新:

发表机构

Uniphore(Uniphore公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过受控消融实验发现,声学与韵律特征组合在话轮结束检测中达到最佳性能(F1=0.93),而添加文本无益,表明话轮转换主要由语调和沉默驱动。

AI 中文摘要

在对话式人工智能中,检测说话者何时结束讲话对于自然的话轮转换至关重要。尽管近期工作引入了语义信息,但不同模态的相对贡献仍不明确。我们提出了一种使用轻量级三模态分类器对声学、韵律和语义信号进行流式话轮结束检测的受控消融研究。在相同的训练条件下,声学与韵律的组合在准确性和延迟之间达到了最佳平衡,实现了话语F1分数0.93,在400毫秒中位延迟下误报率为7.8%。添加文本增加了过早检测,而未提升性能。特征空间分析证实,韵律特征具有最强的类别可分性,而文本表示则存在大量重叠。这些发现表明,话轮转换主要通过语调和沉默模式传达,而非语义完整性,从而无需昂贵的文本推理即可实现更快、更可靠的系统。

英文摘要

In conversational AI, detecting when a speaker has finished talking is crucial for natural turn taking. While recent work incorporates semantics, the relative contribution of different modalities remains unclear. We present a controlled ablation of acoustic, prosodic, and semantic signals for streaming end of turn detection using a lightweight trimodal classifier. Under identical training conditions, the acoustic prosodic combination achieves the best balance of accuracy and latency, achieving utterance F1 of 0.93 with 7.8% false alarms at 400ms median latency. Adding text increases premature detections without improving performance. Feature space analysis confirms that prosodic features have the strongest class separability, while text representations overlap substantially. These findings suggest that turn-taking is primarily conveyed through intonation and silence patterns rather than semantic completeness, enabling faster and more reliable systems without expensive text inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑