arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10765eess.AS

从指标到自然对话:口语对话模型的法语全双工基准

From Metrics to Natural Dialogue: French Full-Duplex Benchmark for Spoken Dialogue Models

Hamid Soltani, Gilles Boulianne

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出法语全双工基准(FDB),含加拿大和欧洲法语变体,评估暂停处理、话轮转换等技能,发现时间指标跨语言稳定,内容指标受语言影响,且指标优化与自然度存在权衡。

中文摘要 AI 辅助

全双工口语对话模型旨在通过允许语音代理在持续对话中倾听、说话、暂停和回应,使其更加自然。然而,当模型在不同语言中评估时,全双工基准是否表现相同尚不清楚。为探究此问题,我们引入了法语全双工基准(FDB),包含两个变体:用于加拿大法语的CALLFC-FDB和用于欧洲法语的MEDIA-FDB,并将它们与英语FDB进行比较。这些基准基于真实口语资源构建,评估关键的全双工技能,包括暂停处理、话轮转换、反馈语和打断。除引入法语FDB外,我们还使用FDB指标评估人-人对话,以更好地理解这些指标在真实对话中的取值。我们的分析揭示,大多数基于时间的指标在不同语言中表现相似,而基于内容的评估在语言不匹配时性能下降。我们还发现优化基准指标与保持对话自然度之间存在权衡。

英文摘要

Full-duplex spoken dialogue models aim to make voice agents more natural by allowing them to listen, speak, pause, and respond during ongoing conversation. However, it is not clear whether full-duplex benchmarks behave the same way when models are evaluated in a different language. To investigate this, we introduce a French full-duplex benchmark (FDB) with two variants, CALLFC-FDB for Canadian French and MEDIA-FDB for European French, and compare them with an English FDB. Built from real spoken resources, these benchmarks evaluate key full-duplex skills, including pause handling, turn-taking, backchannels, and interruptions. Beyond introducing French FDBs, we evaluate human--human conversations with FDB metrics to better understand the values these metrics take in real-world dialogue. Our analysis reveals that most timing-based metrics behave similarly across languages, while content-based evaluation degrades under language mismatch. We also find a trade-off between optimizing benchmark metrics and preserving conversational naturalness.

发表机构

  • Luqia Technologies(Luqia科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑