发表机构
Alibaba Token Foundry(阿里巴巴Token Foundry)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长语音中领域术语转录不准的问题,提出基于LLM的智能体Agentic-GER,利用全局上下文识别并纠正术语,在GigaSpeechBench上显著降低中文语音的B-CER。
AI 中文摘要
语音语言模型的最新进展提高了长音频的自动语音识别(ASR)性能。然而,准确且一致地转录领域特定术语仍然具有挑战性。受大型语言模型(LLM)的世界知识和上下文能力的启发,我们提出了Agentic-GER,一种基于LLM的智能体,用于长语音中的术语纠正。该智能体利用完整转录的全局上下文来识别可疑术语并解决模糊假设。它选择性地重新转录源语音以检查候选纠正,并使用已接受的编辑来指导后续决策。在GigaSpeechBench上使用四种LLM和两种ASR系统进行的实验表明,在中文和英文中,无论是否进行思考,术语性能均得到一致提升。在中文语音上,Agentic-GER相对于Whisper基线,在偏置字符错误率(B-CER)上实现了高达36.8%的相对降低。
英文摘要
Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agentic-GER, an LLM-based agent for terminology correction in long-form speech. The agent uses global context from the full transcript to identify suspicious terms and resolve ambiguous hypotheses. It selectively re-transcribes the source speech to check candidate corrections, and uses accepted edits to guide subsequent decisions. Experiments with four LLMs and two ASR systems on GigaSpeechBench show consistent terminology improvements in both Chinese and English, with and without thinking. On Chinese speech, Agentic-GER achieves up to a 36.8% relative reduction in biased character error rate (B-CER) over the Whisper baseline.
Commentssubmitted to ICASSP 2027