发表机构
Alibaba Token Foundry(阿里通义实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于LLM的ASR,提出多模态对话上下文框架,整合数据流水线、训练与基准,提升实体召回率,并验证历史语音在口音、方言及目标说话人场景中的价值。
AI 中文摘要
对话上下文为自动语音识别(ASR)提供了跨轮次的语义和声学线索,但依赖历史转录文本可能会传播识别错误,并丢弃发音和说话人信息。我们提出了一个用于基于LLM的ASR的多模态对话上下文框架,该框架整合了场景控制的数据流水线、可扩展的多模态上下文训练和系统化评估。我们围绕实体及其易混淆形式构建对话,并将历史用户语音与助手文本回复交错用于监督微调。我们还引入了MM-ContextASR Bench基准,该基准在五个场景中评估上下文理解和实体纠错能力。使用Qwen3-Omni和Step-Audio-2-mini进行的实验揭示了在处理无关和错误历史信息方面的局限性,并表明我们的数据构建和训练提高了上下文利用率,其中多模态上下文在两个模型上均实现了最高的整体实体召回率。进一步针对口音、方言和目标说话人ASR的实验证明了历史语音的价值。基准数据和评估代码可在以下https URL公开获取。
英文摘要
Conversational context provides semantic and acoustic cues across turns for automatic speech recognition (ASR), but relying on historical transcripts can propagate recognition errors and discard pronunciation and speaker information. We present a multimodal conversational-context framework for LLM-based ASR that integrates a scenario-controlled data pipeline, scalable multimodal context training, and systematic evaluation. We construct dialogues around entities and their confusable forms and interleave historical user speech with assistant text responses for supervised fine-tuning. We also introduce MM-ContextASR Bench, which evaluates contextual understanding and entity error correction across five scenarios. Experiments with Qwen3-Omni and Step-Audio-2-mini reveal limitations in handling irrelevant and erroneous history and show that our data construction and training improve context utilization, with multimodal context achieving the highest overall entity recall on both models. Further experiments on accent, dialect, and target-speaker ASR demonstrate the value of historical speech. The benchmark data and evaluation code are publicly available at https://github.com/llh666521/MM-ContextASR.
CommentsTechnical report