arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18132cs.CLcs.SDeess.AS

对齐即全部所需:通用音频-语言模型的无指令训练

Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models

  • Zhejiang University(浙江大学)
  • Tencent Hunyuan(腾讯混元)

机构由 AI 辅助整理,请以论文原文为准。

Xuanru Zhou, Yiwen Shao, Jiahong Li, Dong Yu

AI总结:

该研究提出仅对齐的无指令大音频-语言模型LALM,仅训练轻量投影器,在多模态数据集上用更少数据达到或优于基线,证明仅靠对齐即可构建有竞争力的多模态大语言模型。

AI中文摘要:

多模态大语言模型(MLLMs)通常通过多阶段流程构建,包含跨模态对齐、监督微调(SFT)和偏好优化。该流程假设将大语言模型(LLM)适配至新模态需要大量特定任务监督。然而,预训练LLM已具备强大的推理和指令遵循能力。随着LLM快速发展,一个重要问题仍待解决:我们能否以最少干预将这些能力高效迁移至新模态,且仅靠对齐是否足以构建多模态模型?我们提出仅对齐的无指令大音频-语言模型(LALM),该模型保持音频编码器和LLM完全冻结,仅学习一个轻量投影器。借鉴AzeroS的见解,我们使用自生成数据构建的(音频,响应)对进行训练,其中LLM将字幕扩展为自由形式的响应,无需明确任务指令。在MMAU、MMAR、MMSU和MMAU-Pro数据集上,我们的方法使用少得多的数据,达到或超过经过大量后训练的基线。通过保持LLM冻结,我们的模型保留了其原生指令遵循能力,且可跨模型代际无缝迁移。我们的结果表明,有竞争力的MLLM可仅通过对齐产生,将多模态扩展简化为轻量投影器训练问题,该问题可跨模态泛化并快速适配每个新LLM版本。

英文摘要:

Multimodal large language models (MLLMs) are typically built through a multi-stage pipeline consisting of cross-modal alignment, supervised fine-tuning (SFT), and preference optimization. This pipeline assumes that adapting an LLM to a new modality requires extensive task-specific supervision. However, pretrained LLMs already possess strong reasoning and instruction-following abilities. As LLMs evolve rapidly, an important question remains: can we efficiently transfer these capabilities to a new modality with minimal intervention, and is alignment alone sufficient for building a multimodal model? We introduce an Instruction-Free Alignment-Only large audio-language model (LALM) that keeps both the audio encoder and the LLM fully frozen, learning only a lightweight projector. Borrowing insights from AzeroS [1], we train on (audio, response) pairs from Self-Generated Data Construction, where an LLM expands captions into free-form responses without explicit task instructions. Across MMAU, MMAR, MMSU, and MMAU-Pro, our approach matches or surpasses heavily post-trained baselines using substantially less data. By keeping the LLM frozen, our model preserves its native instruction-following competence and can port seamlessly across model generations. Our results suggest that competitive MLLM can emerge from alignment alone, reducing multimodal extension to a lightweight projector-training problem that generalizes across modalities and adapts rapidly to each new LLM release.

↑