arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DuplexAct-Bench:面向多样化行为需求的主动交互,拓宽全双工语音评估

DuplexAct-Bench: Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements

Keyue Xing, Wentao Ding, Mengmeng Wang, Wenming Tu, Zilong Zheng, Yipeng Kang

arXiv 2609.39446首次发表:更新:

发表机构

State Key Laboratory of General Artificial Intelligence, BIGAI; Peking University; X-LANCE Lab, Shanghai Jiao Tong University(通用人工智能全国重点实验室,北京通用人工智能研究院; 北京大学; 上海交通大学X-LANCE实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有全双工语音基准覆盖不全的问题,提出DuplexAct-Bench双语基准,系统评估六种交互行为,发现当前系统在时序与内容上存在显著差异和失配。

AI 中文摘要

现有的全双工语音基准仅覆盖实时交互行为的子集,且往往局限于有限的上下文条件。我们引入了DuplexAct-Bench,这是一个双语基准,系统地涵盖了六种互补行为,从打断和让位到主动发起、主动沉默和反馈语,并跨越会话前、会话中和无明确条件。在1,290个英汉流式试验中,我们从时序和内容两方面评估了12个全双工语音系统。结果显示,行为、条件和系统之间存在显著差异,并且语义质量与行为时序之间频繁出现不匹配。这些发现表明,当前系统在实时交互展开时,远未能稳健地管理何时、是否以及如何参与。项目页面:此https URL

英文摘要

Existing full-duplex speech benchmarks cover only subsets of real-time interaction behaviors, often under limited contextual conditions. We introduce DuplexAct-Bench, a bilingual benchmark that systematically covers six complementary behaviors, from interruption and yielding to proactive initiation, active silence, and backchanneling, across Pre-session, In-session, and No-explicit conditions. Across 1,290 English and Chinese streaming trials, we evaluate 12 full-duplex speech systems on both Timing and Content. Results reveal substantial variation across behaviors, conditions, and systems, as well as frequent mismatches between semantic quality and behavioral timing. These findings show that current systems remain far from robustly managing when, whether, and how to participate as real-time interaction unfolds. Project page: https://alitaxky.icu/DuplexAct-Bench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑