发表机构
University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究上下文自适应是否为所有编码器-解码器语音识别模型的固有能力,通过六种模型和两种示例形式验证,发现均能实现自适应,且词汇与说话人信息均有贡献,表明该能力具有普遍性。
AI 中文摘要
上下文学习提供了一种有吸引力的方法,通过在推理时提供语音-文本对作为示例,使自动语音识别(ASR)模型适应新的说话人、口音和领域。近期工作表明,当提供交错语音-文本示例时,一些基于大语言模型(LLM)的语音模型能够进行ASR上下文自适应。在本工作中,我们探究上下文自适应是否是所有编码器-解码器模型的固有能力。我们研究了两种示例形式——合并式示例和交错式示例,涵盖六个编码器-解码器模型,包括传统的基于交叉注意力的架构和基于LLM的架构。我们发现,所有测试模型都能开箱即用地进行上下文自适应,在oracle实验中实现了高达30%的相对改进,在使用首轮假设时实现了高达23%的相对改进。通过在三个英文数据集上的受控实验,我们表明词汇信息和说话人信息都对成功的自适应有所贡献。虽然交错式示例在某些情况下有效,但合并式示例在整体上带来一致的自适应效果。我们的结果表明,ASR的上下文自适应并非特定架构、训练或示例方法所独有。
英文摘要
In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, when providing interleaved speech-text demonstrations. In this work, we ask whether in-context adaptation is an inherent ability for all encoder-decoder models. We study two forms of demonstration, collated and interleaved demonstration, across six encoder-decoder models, spanning conventional cross-attention-based and LLM-based architectures. We find that all tested models are able to perform in-context adaptation out of the box, achieving up to 30% relative improvement in the oracle experiments and up to 23% using first-pass hypotheses. Through controlled experiments on three English datasets, we show that lexical and speaker information both contribute to successful adaptation. While interleaved demonstration is effective in certain cases, collated demonstration brings consistent adaptation across the board. Our results suggest that in-context adaptation for ASR is not unique to specific architectures, training, or demonstration approaches.
CommentsAccepted to IEEE SLT 2026