arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TurnBench:面向口语对话中话轮转换动态的多领域基准测试

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe

arXiv 2608.25218首次发表:更新:

发表机构

Sesame AI; Mundo AI; Carnegie Mellon University; National Taiwan University; Academia Sinica; Oto; Brno University of Technology(Sesame AI; Mundo AI; 卡内基梅隆大学; 台湾大学; 中央研究院; Oto; 布尔诺理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对口语对话话轮转换评估的局限,提出多领域基准测试TurnBench,含标注语料与评估协议,测试14种系统发现打断误报具风格依赖性,当前系统未达人类流畅交接表现。

AI 中文摘要

自然对话中的说话者会轮流发言与倾听,实时决定何时接过、保持或让出话语权。然而,由于缺乏一致的、基于语言学的评估协议以及覆盖多样对话类型的人工标注数据,话轮转换评估仍存在局限。为解决该问题,我们提出TurnBench,这一多领域基准测试将时长30小时的人工标注双人对话语料库与标准化的话轮结束及打断检测评估协议相结合。我们将对话类型设为可控实验变量,涵盖6种不同的交互风格,并对每段对话进行三重标注。对14种异构话轮转换系统进行基准测试后,我们发现话轮结束召回率在各类风格中保持稳定,而打断误报则强烈依赖于对话类型,且集中在 backchannel(反馈通道)密集的交互风格中。尽管在流畅的话语权交接中,人类说话者平均会在当前话轮结束前151毫秒开始发言,但目前没有任何系统能达到同等表现且不会产生过多误报。我们在此httpsURL发布了该语料库、一个时长104小时的训练集以及带有交互式数据集查看器的公开排行榜。

英文摘要

Speakers in natural conversation take turns speaking and listening, deciding in real time when to take, hold, or yield the floor. However, turn-taking evaluation remains limited due to the lack of a consistent, linguistically grounded evaluation protocol and hand-annotated data covering diverse conversation types. To address this, we present TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection. We set conversation type as a controllable experimental variable, covering six distinct interaction styles, and triple-annotate each conversation. Benchmarking 14 heterogeneous turn-taking systems, we find end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in backchannel-dense interaction styles. Although in smooth floor transfers human listeners begin speaking a median 151 ms before the current turn ends, no current system performs equivalently without incurring excessive false positives. We release our corpus, a 104-hour training set, and a public leaderboard with an interactive dataset viewer at https://turnbench.sesame.com.

Comments8 pages, 2 figures. Accepted to IEEE SLT 2026. v2: camera-ready version

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑