arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

教LLM分辨谁说了什么:面向无编码器语音LLM的元数据监督预训练

Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs

Mohan Shi, Ruchao Fan, Sunit Sivasankaran, Keqi Deng, Jinyu Li

arXiv 2610.01695首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无编码器语音LLM预训练不足的问题,提出元数据监督预训练(MSP)及说话人感知话语组合(SAUC),在多方对话ASR和说话人日志任务上优于随机初始化的编码器模型,并可与大规模预训练编码器模型竞争。

AI 中文摘要

基于编码器的语音大语言模型(Speech-LLMs)通常采用预训练语音编码器,这些编码器优先考虑语言内容,但可能丢弃对说话人区分和副语言理解至关重要的细粒度声学线索。无编码器语音大语言模型则通过轻量级嵌入层将梅尔频谱图特征直接映射到LLM输入空间,使LLM能够从低级声学特征中学习。然而,针对无编码器语音大语言模型的系统性预训练策略仍未得到充分探索,限制了它们弥补缺乏大规模预训练语音编码器这一不足的能力。我们提出了元数据监督预训练(MSP),利用说话人身份和情感等语音属性来发展说话人区分和副语言能力。我们进一步引入了说话人感知话语组合(SAUC)以加强说话人区分,并应用随机跨度掩码来规范化预训练。我们主要在多方对话中的联合自动语音识别(ASR)和说话人日志任务上评估我们的方法,并辅以副语言语音理解任务的实验。在匹配的训练数据条件下,我们的无编码器模型优于随机初始化的基于编码器的对应模型。在元数据标注数据有限的情况下,它与使用在更大规模语料库上预训练的语音编码器的模型具有竞争力,并在多种设置中超越它们。这些结果证明了无编码器架构在构建获得多样化语音能力的原生多模态LLM方面的潜力。

英文摘要

Encoder-based speech large language models (Speech-LLMs) commonly employ pretrained speech encoders that prioritize linguistic content but may discard fine-grained acoustic cues essential for speaker discrimination and paralinguistic understanding. Encoder-free Speech-LLMs instead map Mel-spectrogram features directly into the LLM input space through lightweight embedding layers, enabling the LLM to learn from low-level acoustic features. However, systematic pretraining strategies for encoder-free Speech-LLMs remain underexplored, limiting their ability to compensate for the absence of large-scale pretrained speech encoders. We propose metadata-supervised pretraining (MSP), which leverages speech attributes such as speaker identity and emotion to develop speaker-discriminative and paralinguistic capabilities. We further introduce speaker-aware utterance composition (SAUC) to strengthen speaker discrimination and apply random span masking to regularize pretraining. We primarily evaluate our approach on joint ASR and speaker diarization in multi-speaker conversations, complemented by experiments on paralinguistic speech-understanding tasks. Under matched training-data conditions, our encoder-free model outperforms its randomly initialized encoder-based counterpart. With limited metadata-annotated data, it is competitive with models using speech encoders pretrained on substantially larger corpora, outperforming them in several settings. These results demonstrate the potential of encoder-free architectures for building native multimodal LLMs that acquire diverse speech capabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑