CoSE-E:企业场景下的语码转换语音评估基准
CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
浏览论文内容
中文总结 AI 辅助
本文提出CoSE-E基准,用于企业场景下语码转换语音识别评估,涵盖多维度框架、5种语言对的系统评测及错误诊断分析,以支持多语言语音代理部署。
中文摘要 AI 辅助
语码转换(CS)指在同一话语中无缝切换语言,这仍是自动语音识别(ASR)中的关键挑战。尽管先前工作聚焦于对话式CS-ASR,企业场景要求评估超出编辑距离误差的操作影响:语码转换转录错误如何传播到下游语音代理任务失败。在这项工作中,我们提出(1)一个面向企业领域的CS-ASR合成基准和多维度评估框架,(2)对前沿ASR系统在5种语言对上的系统评估,(3)对语码转换在不同语言对和模型中引入的额外转录错误的诊断分析。我们发布CoSE-E以支持企业部署中多语言语音代理的企业级CS-ASR评估。
英文摘要
Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how code-switching transcription errors propagate to downstream voice agent task failures. In this work, we propose (1) a CS-ASR synthetic benchmark and multidimensional evaluation framework tailored to enterprise domains, (2) systematic evaluation of frontier ASR systems across 5 language pairs, (3) diagnostic analysis of the additional transcription errors that code-switching introduces across language pairs and models. We release COSE-E to support enterprise-focused CSASR evaluation for multilingual voice agents in enterprise deployment.
发表机构
- ServiceNow AI Research(ServiceNow AI研究院)
- Qualcomm Technologies, Inc.(高通技术公司)
机构由 AI 辅助整理,请以论文原文为准。