arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

共生架构:冻结语言模型的后期音频扩展

Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models

Yotaro Kubo, Qi Sun, Yujin Tang

arXiv 2609.30784首次发表:更新:

发表机构

Sakana AI(Sakana AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种共生架构,通过注入器向冻结LLM的KV缓存写入音频向量,实现无需微调的音频理解,兼具可扩展性与性能,接近微调ALM且保留文本能力。

AI 中文摘要

本文提出了一种架构,用于在不微调权重的情况下,为大型语言模型(LLM)配备音频理解能力。所提出的共生架构采用了一个注入器模块,该模块将音频条件向量直接写入目标LLM的短期记忆,即键值(KV)缓存,从而使LLM能够表现为音频语言模型(ALM)。该架构的优势有两方面。首先,它提高了ALM的可扩展性:由于所提方法在音频注入期间绕过了LLM,注入成本由注入器宽度而非主干网络宽度决定,因此其扩展速度可以比全主干预填充的成本更慢。其次,由于训练方案不更新LLM权重,LLM的原始能力得以保留,不存在因微调而退化的风险。所提方法的有效性在音频理解任务(自动语音识别、音频问答和声学场景分类)以及纯文本任务上进行了评估。我们确认,在音频预填充期间激活更少参数的情况下,我们的架构优于使用冻结LLM的传统方法,并接近微调ALM的性能,同时通过构造保留了主干LLM原有的纯文本任务性能。

英文摘要

This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memory, i.e., the key-value (KV) cache, enabling the LLM to behave as an audio language model (ALM). The architectural advantages are twofold. First, it improves the scalability of ALMs: because the proposed method bypasses the LLM during audio injection, the injection cost is governed by the injector width rather than the backbone width, and can therefore scale more slowly than the cost of full-backbone prefilling. Second, since the training scheme does not update the LLM weights, the original capabilities of the LLM are preserved without the risk of degradation from fine-tuning. The effectiveness of the proposed method is evaluated on both audio-understanding tasks (automatic speech recognition, audio question answering, and acoustic scene classification) and text-only tasks. We confirm that, while activating fewer parameters during audio prefilling, our architecture outperforms the conventional method with a frozen LLM and approaches the performance of a fine-tuned ALM, all while preserving the backbone LLM's original text-only task performance by construction.

CommentsSubmitted to ICASSP

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑