arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35820cs.CLcs.AIcs.SDeess.AS

$τ$-Multilingual:跨语言语音代理基准测试

$τ$-Multilingual: Benchmarking Voice Agents Across Languages

Soham Ray, Edgard dos Santos Paiva, Ruben Valenzuela, Karthik Narasimhan, Keshav Dhandhania, Victor Barres

首次发表
浏览论文内容

中文总结 AI 辅助

提出τ-Multilingual多语言语音代理基准,扩展至五种语言,发现韩语和普通话任务完成度显著下降,并发布语言包与评估工具。

中文摘要 AI 辅助

仅使用英语的基准测试只能展现语音代理行为的一小部分。我们推出了 $\ au$-Multilingual,将 $\ au$-Voice 扩展到西班牙语、巴西葡萄牙语、印地语、韩语和普通话,并通过母语者审查和评估生成的语音及口语输出。在 4,500 次全双工通话和五种语音配置中,西班牙语、葡萄牙语和印地语的任务完成度与英语相差在 3.2 个百分点以内,但韩语和普通话分别下降了 14.7 和 8.4 个百分点。失败模式也各不相同:韩语系统漏答更多,普通话系统打断更频繁,两者在工具和实体处理上均存在困难。Grok 在任务完成度上领先,但在生成质量上得分最低,这促使我们分别报告任务、交互和生成指标。我们发布了语言包、经过验证的评估器以及用于社区构建多语言语音代理评估的工具。

英文摘要

English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce $τ$-Multilingual, extending $τ$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output. Across 4,500 full-duplex calls and five voice configurations, Spanish, Portuguese, and Hindi remain within 3.2 task-completion points of English, but Korean and Mandarin fall by 14.7 and 8.4 points. The failure modes also vary: Korean systems miss more responses, Mandarin systems interrupt more often, and both struggle with tools and entities. Grok leads task completion but scores lowest on generation quality, motivating separate task, interaction, and generation reporting. We release language packs, validated judges, and tools for community-built multilingual voice-agent evaluation.

发表机构

  • Sierra
  • Mercor
  • Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑