发表机构
NatWest(国民西敏寺银行)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语音金融助手中英国地区口音识别鲁棒性不足的问题,本文提出首个内部基准CavaBench,评估ASR-LLM流水线,发现WER强预测工具调用准确性但不足以反映任务性能,为设计包容可靠的金融语音助手提供指导。
AI 中文摘要
AI语音助手通常使用自动语音识别(ASR)结合基于大语言模型(LLM)的推理,然而现有系统在处理英国地区口音(包括苏格兰口音、爱尔兰口音和威尔士口音)时表现不佳,因为大多数ASR模型主要使用美国英语语音数据进行训练。因此,错误可能传递到LLM阶段,破坏工具调用参数并产生错误或缺失的响应,这在金融领域尤其代价高昂。可部署的ASR还必须满足严格的延迟和内存预算,这使得选择对口音鲁棒的模型更加困难。我们引入了CavaBench,这是首个内部收集的口语金融查询基准,并利用它评估了一系列ASR模型及其端到端ASR-LLM流水线在不同自我报告的英国口音下的表现。我们发现词错误率(WER)能强预测下游工具调用准确性(r = -0.93),但可能无法反映任务级性能,且与口音相关的失败在不同模型和声学条件下差异显著。这些发现为设计更具包容性、更可靠的语音金融助手提供了指导。
英文摘要
AI voice assistants often use Automatic Speech Recognition (ASR) with LLM-based reasoning, yet existing systems struggle with regional British accents, including Scottish, Irish, and Welsh accents, since most ASR models are trained predominantly on American English voice data. Consequently, errors can carry through to the LLM stage, corrupting tool-call arguments and producing wrong or missing responses, which is especially costly in finance. Deployable ASR must also meet tight latency and memory budgets, making an accent-robust model choice even harder. We introduce CavaBench, the first internally collected benchmark of spoken financial queries, and use it to evaluate a range of ASR models and their end-to-end ASR-LLM pipeline behaviour across self-reported British accents. We find that WER strongly predicts downstream tool-calling accuracy ($r = -0.93$) but can fail to reflect task-level performance, with accent-related failures varying substantially across models and acoustic conditions. These findings guide the design of more inclusive, reliable voice-based financial assistants.
CommentsICASSP 2027 submission