发表机构
AI Center, OPPO(OPPO AI中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HearInContext基准通过同音词测试隐式上下文,微调Qwen3-ASR-1.7B显著提升目标词召回率且不损害通用性能。
AI 中文摘要
上下文感知的语音识别(Contextual ASR)可以从语义线索或上下文中明确提供的目标词中获益。我们引入了HearInContext,一个中英双语基准,它将共享的合成语音与支持不同解释的助手回复配对。该基准包含3,764个围绕同音词构建的语义测试用例。隐式上下文排除候选词;显式上下文则命名目标词。无上下文和无关上下文对照组用于衡量相关历史信息的益处以及对无关历史信息的敏感性。具备上下文能力的模型能从隐式线索中获益,但在显式提示下能实现更高的目标词召回率。对Qwen3-ASR-1.7B进行微调后,在中文和英文中,隐式上下文的目标词召回率分别提升了11.0和11.5个百分点,而在AISHELL-1和LibriSpeech上的绝对CER/WER变化保持在0.1个百分点以下。这些提升也扩展到了微调中未包含的显式条件,以及真实录音上的中文热词识别。
英文摘要
Contextual ASR can benefit from semantic cues or from target words explicitly provided in the context. We introduce HearInContext, a Mandarin-English benchmark that pairs shared synthetic speech with assistant replies supporting different interpretations. The benchmark comprises 3,764 semantic test cases built around homophones. Implicit contexts exclude candidate words; explicit contexts name the target. No-context and unrelated-context controls measure the benefit of relevant history and sensitivity to irrelevant history. Context-capable models benefit from implicit cues but achieve higher target recall with explicit hints. Fine-tuning Qwen3-ASR-1.7B improves implicit-context target recall by 11.4 percentage points in both Mandarin and English, while absolute CER/WER changes on AISHELL-1 and LibriSpeech remain below 0.1 percentage points. Gains extend to explicit conditions excluded from fine-tuning and to Mandarin hotword recognition on real recordings. Code and data are available at https://github.com/OPPO-Mente-Lab/HearInContext
CommentsSubmitted to ICASSP 2027