发表机构
University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低资源语音识别,提出基于Fisher白化互协方差的任务信息感知适配方法,通过自适应秩分配在固定参数预算下提升Whisper和Qwen3-ASR的性能,优于LoRA。
AI 中文摘要
将多语言语音基础模型适配到低资源语言仍然困难,尤其是对于在预训练期间代表性不足的语言。虽然参数高效微调(PEFT)降低了适配大型模型的成本,但诸如LoRA之类的传统方法依赖于通用的低秩参数化,并未显式使用下游任务信息来定义适配子空间。为了研究任务信息感知的PEFT是否能更好地支持低资源自动语音识别(ASR),我们将Fisher白化互协方差分析(FCCA)应用于Whisper和Qwen3-ASR,并引入了两个互补的扩展:非对称耦合FCCA(AC-FCCA),其利用结构化的跨层共享;以及自适应秩FCCA(AR-FCCA),其在固定参数预算下跨投影矩阵重新分配适配容量。在受控的多语言实验中,我们在预训练期间代表性不足或不受支持的语言以及代表性良好的语言上评估了这些方法。标准FCCA与可训练参数预算匹配的LoRA相比具有竞争力,并且通常优于后者。AR-FCCA在两种模型架构上相对于标准FCCA提供了最一致的改进,在多个评估设置中具有统计学显著提升,同时保持相同数量的可训练参数。这些结果表明,任务信息感知的子空间构建对于低资源语音适配是有效的,并且自适应秩分配提供了一种在不增加模型容量的情况下提高参数效率的稳健方法。
英文摘要
Adapting multilingual speech foundation models to low-resource languages remains difficult, especially for languages that are poorly represented during pre-training. While parameter-efficient fine-tuning (PEFT) reduces the cost of adapting large models, conventional approaches such as LoRA rely on generic low-rank parameterizations and do not explicitly use downstream task information to define the adaptation subspace. To investigate whether task-informed PEFT can better support low-resource ASR, we apply Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR, and introduce two complementary extensions: Asymmetric-Coupled FCCA (AC-FCCA), which exploits structured cross-layer sharing, and Adaptive-Rank FCCA (AR-FCCA), which reallocates adaptation capacity across projection matrices under a fixed parameter budget. Under controlled multilingual experiments, we evaluate these approaches on languages that are poorly represented or unsupported during pre-training alongside well-represented languages. Standard FCCA is competitive with, and usually outperforms, trainable-parameter-budget-matched LoRA. AR-FCCA provides the most consistent improvement over standard FCCA across both model architectures, with statistically significant gains in several evaluation settings, while retaining the same number of trainable parameters. These results show that task-informed subspace construction can be effective for low-resource speech adaptation, and that adaptive rank allocation provides a robust way to improve parameter efficiency without increasing model capacity.