利用韩语语音识别中的细粒度错误纠正提升咨询服务
Leveraging Fine-grained Error Correction in Korean Speech Recognition for Consultation Services
浏览论文内容
中文总结 AI 辅助
针对韩语呼叫中心ASR转录错误,提出首个大规模基准DasanCallDial和检测器门控上下文跨度纠正(DCSC)框架,通过多级粒度检测与纠正实现最先进性能。
中文摘要 AI 辅助
自动语音识别(ASR)技术是客户服务自动化和大规模转录的基础。然而,即使在复杂的现实环境(如呼叫中心对话)中,先进的ASR模型也会出现不可避免的错误。当隐私限制排除音频访问时,错误纠正必须依赖于基于文本的后期编辑。现有的纯文本方法在低资源语言中面临重大挑战,主要原因是标注语料库和定制化纠正方法的严重稀缺。对于韩语而言,这种资源缺口尤为突出,因为现有资源主要设计用于ASR训练而非基于文本的错误纠正。为解决这一问题,我们引入了DasanCallDial,这是首个专门为对话级ASR错误纠正而策划的大规模韩语基准数据集。该数据集源自真实的呼叫中心交互,包含1,974个对话和115,460个话语。利用这一资源,我们提出了检测器门控上下文跨度纠正(DCSC),这是一种用于错误稀疏的韩语语音识别转录本的纯文本后期编辑框架。DCSC结合了基于编码器的检测器,该检测器首先执行词元级错误检测,随后由基于语言模型的纠正器进行训练,以纠正细粒度的跨度级错误。此外,我们采用对话级上下文增强,使模型能够利用话语历史进行消歧。通过采用多级粒度,我们的方法实现了最先进的性能,有效克服了通用大语言模型在低资源环境中的局限性。
英文摘要
Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR models exhibit inevitable errors in complex real-world environments such as call center conversations. When privacy restrictions preclude audio access, error correction must rely on text-based post-editing. Existing text-only approaches face significant challenges in low-resource languages, mainly due to a critical scarcity of annotated corpora and tailored correction methodologies. For Korean, this resource gap is particularly pronounced, as existing resources are predominantly designed for ASR training rather than text-based error correction. To address this, we introduce DasanCallDial, the first large-scale Korean benchmark dataset specifically curated for dialogue-level ASR error correction. Derived from genuine call center interactions, it comprises 1,974 dialogues with 115,460 utterances. Leveraging this resource, we propose Detector-Gated Contextual Span Correction (DCSC), a text-only post-editing framework for error-sparse Korean speech recognition transcripts. DCSC combines an encoder-based detector that first performs token-level error detection, followed by a language model-based corrector trained to rectify fine-grained span-level errors. Additionally, we employ dialogue-level context augmentation to enable the model to leverage discourse history for disambiguation. By employing multi-level granularity, our method achieves state-of-the-art performance, effectively overcoming the limitations of general LLMs in low-resource settings.
发表机构
- Chung-Ang University(中央大学)
- Korea Local Information Research & Development Institute(韩国地方信息研究院)
- SK intellix
机构由 AI 辅助整理,请以论文原文为准。