DirectSpeech2LLM:一种缓解语音大语言模型中提示过拟合的简单端到端框架
DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs
浏览论文内容
中文总结 AI 辅助
DirectSpeech2LLM提出一种简单端到端框架,通过基于距离的CTC损失对齐语音嵌入,缓解语音LLM的提示过拟合,实现零样本泛化至语音翻译和情感识别。
中文摘要 AI 辅助
语音大语言模型(Speech-LLMs)常常表现出提示过拟合现象,即仅使用自动语音识别(ASR)指令训练的模型难以泛化到诸如语音翻译等新指令上,并持续主要作为ASR系统运行。我们提出了DirectSpeech2LLM,一个简单的端到端框架,在语音条件下保留了LLM在未见任务上的指令遵循能力。该框架在冻结的LLM嵌入矩阵上计算基于距离的CTC损失,并分别使用贪婪CTC标签来推导几何和时间对齐的语音嵌入作为LLM的输入。仅使用960小时的LibriSpeech ASR数据进行训练,DirectSpeech2LLM在ASR(已见任务)上优于级联系统,并零样本泛化到语音翻译和情感识别(两个未见任务),在这两个新指令上接近级联系统的上界,尽管训练期间从未见过它们。我们还发现,几何对齐强度所起的作用比先前假设的要小,因为我们的修改版CTC损失被证明提供了足够的隐式几何基础,无需显式回归损失。结果在两个LLM家族中保持一致,并随着更多训练数据和模型容量的增加而扩展。
英文摘要
Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to speech translation and emotion recognition (two unseen tasks), closely matching the cascaded system upper bound on these two new instructions despite seeing neither during training. We also find that geometric alignment strength plays a smaller role than previously assumed, as our modified CTC loss is shown to provide sufficient implicit geometric grounding without requiring an explicit regression loss. Results are consistent across two LLM families and scale with both more training data and model capacity.
发表机构
- IIIT Delhi(德里印度理工学院)
- Microsoft Research(微软研究院)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。