通过适配器唤醒编码器:语音大语言模型的有效域自适应微调
Encoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs
浏览论文内容
中文总结 AI 辅助
针对语音大语言模型在域偏移语音(如儿童和方言)上适应困难的问题,提出EAVA方法,通过在编码器层插入适配器并联合微调,显著提升ASR性能并达到新最优。
中文摘要 AI 辅助
语音大语言模型(Speech-LLMs)通常由预训练的语音编码器、模态投影器以及使用低秩适配器(LoRA)微调的大语言模型(LLM)构成,在通用领域语音上展现出强大的自动语音识别(ASR)性能。然而,在目标域数据有限的情况下,将其适配到域偏移语音(如儿童语音或方言语音)仍然具有挑战性。鉴于LLM在语音大语言模型中的主导作用,且交叉熵损失仅应用于LLM输出,语音编码器可能无法充分适应新的声学条件。在本文中,我们提出了通过适配器唤醒编码器(EAVA),一种简单而有效的基于语音大语言模型的ASR域自适应微调方法。首先,将轻量级适配器插入每个编码器层并仅对其进行训练,使目标域声学知识能够融入编码器,同时保留其预训练知识。其次,在目标域上对完整模型进行联合微调,其中LLM应用LoRA。在三个域偏移ASR数据集(涵盖儿童语音和方言语音)上的实验表明,EAVA始终优于普通微调和其他基线方法,取得了新的最先进性能。
英文摘要
Speech Large Language Models (Speech-LLMs), typically built from a pre-trained speech encoder, a modality projector, and an LLM fine-tuned with Low-Rank Adapters (LoRA), have shown strong Automatic Speech Recognition (ASR) performance on general-domain speech. However, adapting them to domain-shifted speech, such as child or dialectal speech, remains challenging under limited target-domain data. Given the dominant role of the LLM in Speech-LLMs, with cross-entropy loss applied only at the LLM output, the speech encoder may receive insufficient adaptation to new acoustic conditions. In this paper, we propose Encoder Awakening via Adapters (EAVA), a simple yet effective domain-adaptive fine-tuning method for Speech-LLM-based ASR. First, lightweight adapters are inserted into each encoder layer and trained exclusively, enabling target-domain acoustic knowledge to be incorporated into the encoder while preserving its pre-trained knowledge. Second, the full model is jointly fine-tuned on the target domain with LoRA applied to the LLM. Experiments on three domain-shifted ASR datasets, covering child and dialectal speech, show that EAVA consistently outperforms vanilla fine-tuning and other baselines, achieving new state-of-the-art performance.
发表机构
- Department of Electrical and Computer Engineering University of California Los Angeles(加利福尼亚大学洛杉矶分校电气与计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。