arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越WER:带口音对话式ASR中的实体与不流畅召回

Beyond WER: Entity and Disfluency Recall in Accented Conversational ASR

Fiza Husain, Ankit Pandey, Yash Singh

arXiv 2609.20828首次发表:更新:

发表机构

Stimuler(Stimuler)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对带口音对话式ASR,提出三阶段流水线,通过数据策展和区域LoRA适配器,将实体召回率提升至80-85%,填充词召回率提升至76-86%,并保持低WER。

AI 中文摘要

针对词错误率(WER)优化的ASR系统常常遗漏带口音对话式英语中的命名实体和填充停顿,而这两者对于语言学习反馈至关重要。我们针对来自印度、印度尼西亚和拉丁美洲的说话者提出了一种三阶段流水线:(1)启发式SQL过滤器,以随机采样2.8倍的实体密度策展富含实体的训练数据;(2)在Qwen2.5-Omni-3B上微调的区域LoRA适配器,在单次前向传播中同时生成逐字和纠正后的转录;(3)由基于LLM的评判者验证的六类别错误分类法(83.8%一致性,210个人工标注样本)。该流水线在6k测试话语上实现了80-85%的实体召回率(从53-55%提升)、76-86%的填充词召回率(从<5%提升)和6-10%的WER,在实体召回率上优于Whisper和商业ASR,同时以少10倍的参数匹配零样本30B模型。配对自助法检验确认,仅策展就贡献了2.8-4.2个百分点的实体召回率提升(p<0.0001)。

英文摘要

ASR systems optimised for Word Error Rate (WER) often miss named entities and filled pauses in accented conversational English, both critical for language-learning feedback. We present a three-stage pipeline for speakers from India, Indonesia, and Latin America: (1) heuristic SQL filters curating entity-rich training data at 2.8x the entity density of random sampling, (2) regional LoRA adapters fine-tuned on Qwen2.5-Omni-3B producing both verbatim and corrected transcripts in a single forward pass, and (3) a six-category error taxonomy validated by an LLM-based judge (83.8% agreement, 210 human-labelled samples). The pipeline achieves 80-85% entity recall (up from 53-55%), 76-86% filler recall (up from <5%), and 6-10% WER across 6k test utterances, outperforming Whisper and a commercial ASR on entity recall while matching a zero-shot 30B model with 10x fewer parameters. Paired bootstrap tests confirm that curation alone accounts for 2.8-4.2 pp of entity recall gain (p<0.0001).

Comments5 pages, 1 figure, 1 table, accepted at Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑