arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18516cs.CL

对齐、整合与触发:面向零样本语音大语言模型的高效令牌级对齐

Align, Integrate, and Fire: Efficient Token-Level Alignment for Zero-Shot SpeechLLMs

Abderrahmane Issam, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出对齐连续整合与触发框架,通过动态时间规整将声学帧压缩为文本令牌长度,结合轻量级距离度量与单层知识蒸馏,实现零样本语音处理的高效性能。

中文摘要 AI 辅助

虽然大语言模型在自然语言处理方面表现出色,但高效地将其能力扩展到语音输入仍然是一个重大挑战。现有的构建语音大语言模型的方法通常依赖于计算成本高昂的全模型微调,或采用参数高效的投影器,但这些方法存在令牌序列长度低效和全模型监督成本高昂的问题。在本文中,我们引入了对齐连续整合与触发(Aligned Continuous Integrate-and-Fire),这是一个用于零样本语音处理的高效框架。我们的方法利用显式动态时间规整对齐,将连续的声学帧动态压缩为目标文本的精确离散令牌长度。这使得我们的初始训练阶段能够使用轻量级距离度量建立稳健的声学到语义桥梁,完全绕过了计算成本高昂的大语言模型前向传播。对于后续微调,我们提出了一种内存高效的知识蒸馏目标,该目标针对单个大语言模型层,以一小部分计算成本实现了与全模型交叉熵训练相当的性能。通过在自动语音识别和语音翻译上的广泛评估,我们证明了我们的方法相比之前的参数高效基线取得了更优的性能。

英文摘要

While Large Language Models excel in natural language processing, efficiently extending their capabilities to spoken input remains a significant challenge. Existing methods for building SpeechLLMs often rely on computationally expensive full-model fine-tuning, or employ parameter-efficient projectors that suffer from inefficient token sequence lengths and costly full-model supervision. In this paper, we introduce Aligned Continuous Integrate-and-Fire, a highly efficient framework for zero-shot speech processing. Our method dynamically compresses continuous acoustic frames into the exact discrete token length of the target text utilizing explicit Dynamic Time Warping alignments. This allows our initial training stage to establish a robust acoustic-to-semantic bridge using lightweight distance metrics, entirely bypassing the computationally expensive LLM forward pass. For subsequent fine-tuning, we propose a memory-efficient knowledge distillation objective that targets a single LLM layer, performing competitively with full-model cross-entropy training at a fraction of the computational cost. Through extensive evaluations on Automatic Speech Recognition and Speech Translation, we demonstrate that our method achieves superior performance compared to prior parameter-efficient baselines.

发表机构

  • Maastricht University(马斯特里赫特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑