arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

mu-bench:多语言话语转录基准

mu-bench: A Multilingual Utterance Transcription Benchmark

Andrea Li, Soham Ray

arXiv 2609.32082首次发表:更新:

AI 中文总结

针对ASR评估忽视语义的问题,提出mu-bench多语言话语转录基准及UER指标,以LLM评判语义保留,在1,847条人工评分转录上κ=0.78,并排名六家商业ASR提供商,最佳UER为11.9%,普通话最难。

AI 中文摘要

语音代理依赖于准确的自动语音识别(ASR)来根据呼叫者所说内容采取行动,然而ASR是在以英语为中心的朗读语音上评估的,使用词错误率(WER)作为指标,该指标惩罚表面差异而非语义差异。我们引入了mu-bench,这是一个包含250个电话通话中4,270条呼叫者话语的数据集,这些通话涉及一个AI银行代理,使用英语、西班牙语、土耳其语、越南语和普通话,聚焦于表单字段输入,如姓名、电子邮件地址和确认码。我们发布了话语错误率(UER),这是一个LLM评判器,用于判断转录是否保留语义,并与人类评分者进行校准,同时发布一个LLM归一化器,使WER在不同提供商的输出格式之间具有可比性。在1,847条人工评分的转录上,UER与标注者的一致性为κ=0.78,而归一化文本上的精确匹配WER的一致性为0.53。我们在一个公开排行榜上对六家商业提供商进行排名;最佳提供商达到11.9%的UER,而普通话对全部六家提供商来说都是最难的。

英文摘要

Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an AI banking agent in English, Spanish, Turkish, Vietnamese, and Mandarin, centered on form-field inputs such as names, email addresses, and confirmation codes. We release Utterance Error Rate (UER), an LLM judge of whether a transcript preserves meaning, calibrated against human raters, together with an LLM normalizer that makes WER comparable across providers' output formats. On 1,847 human-rated transcripts, UER agrees with annotators at $κ$ = 0.78, versus 0.53 for exact-match WER on normalized text. We rank six commercial providers on a public leaderboard; the best reaches 11.9% UER, and Mandarin is hardest for all six.

Comments5 pages, 7 tables. Dataset: https://huggingface.co/datasets/sierra-research/mu-bench . Code: https://github.com/sierra-research/mu-bench/tree/icassp-2027 . Leaderboard: https://research.sierra.ai/mubench/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑