发表机构
KU Leuven; Academia Sinica(鲁汶大学; 中央研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对语音-LLM桥接模块,提出局部原型重构(LPR)正则化器,通过量化词法流形兼容性提升其迁移性能,在多语种ASR和语音翻译任务上效果显著,且诊断方法可预测桥接模块的可复用性。
AI 中文摘要
语音-大语言模型(LLM)系统通常通过一个小型可训练桥接模块,将冻结的语音编码器与冻结的大语言模型连接起来。该桥接模块常被视为“管道”,但实际上它定义了语音-LLM接口的几何结构,而预训练目标决定了该接口是否能为下游任务提供可复用的初始化。我们通过两个互补属性研究可迁移的桥接模块:与文本侧的全局对齐,以及局部词法流形兼容性——即桥接嵌入保持在冻结LLM输入嵌入的邻域附近。我们提出一种固定的、无头部、无时间戳的诊断方法,可适用于任何预训练目标,以此将该属性量化,并证明下一词预测(NWP)和句子级对比预训练无法完全捕捉词级词法兼容性。随后我们引入局部原型重构(LPR),这是一种轻量级的仅训练阶段正则化器,要求每个对齐的桥接词元可从冻结LLM词元嵌入的小邻域中重构,其极限情况为硬单原型锚点。在多语种自动语音识别(ASR)和语音翻译任务上,LPR提升了迁移性能,在翻译和低资源适应任务上获得最大增益。关键的是,我们的独立诊断方法与各目标下的下游增益相关,表明词法流形兼容性可预测语音-LLM桥接模块的可复用性。
英文摘要
Speech-to-LLM systems often connect a frozen speech encoder to a frozen large language model (LLM) through a small trainable bridge. The bridge is usually treated as plumbing, but it in fact defines the geometry of the speech-to-LLM interface, and the pretraining objective decides whether that interface provides a reusable initialization for downstream tasks. We study a transferable bridge through two complementary properties: global alignment with the text side, and local lexical manifold compatibility, where bridge embeddings remain close to the frozen LLM's input-embedding neighbourhoods. We make this property measurable with a fixed, head-free, timestamp-free diagnostic that applies to any objective, and show that next-word prediction (NWP) and sentence-level contrastive pretraining do not fully capture token-level lexical compatibility. We then introduce Local Prototype Reconstruction (LPR), a lightweight training-only regularizer that requires each aligned bridge token to be reconstructable from a small neighbourhood of frozen LLM token embeddings, with a hard single-prototype anchor as its limiting case. On multilingual ASR and speech translation, LPR improves transfer, with the largest gains on translation and low-resource adaptation. Crucially, our independent diagnostic correlates with downstream gains across objectives, suggesting that lexical manifold compatibility is predictive of reusability for speech-to-LLM bridges.