Listen-to-Reason:借助专家聆听、基于图检索、用大语言模型推理
Listen-to-Reason: Listen with Experts, Retrieve over a Graph, Reason with LLMs
浏览论文内容
中文总结 AI 辅助
本文提出可解释的Listen-to-Reason流水线,用冻结专家编码器映射音频到语义节点、纯文本LLM推理,在SAKURA等数据集上表现优异,新领域适配成本低且可通过强阅读器缩小性能差距。
中文摘要 AI 辅助
大型音频语言模型(LALM)通过多阶段训练将音频编码器融合进大型语言模型(LLM)。这种耦合意味着新领域或更强的LLM需要重新训练,且其答案无法追溯到模型所听内容:思维链是事后解释。我们提出Listen-to-Reason(L2R),一种设计层面即可解释的流水线,它通过显式、人类可读的树将音频传递给LLM:冻结的专家编码器上的小型头将音频片段的每个块映射到树上语义有意义的节点(适用于语音、音乐和环境声音),而冻结的纯文本LLM仅从这些节点和ASR转录本中回答,无需聆听音频片段。因此每个答案都可追溯到它所读取的节点和转录本,且这些节点具有因果性:在SAKURA数据集上,将决定节点替换为干扰项会推翻78%的正确答案。使用7B规模的阅读器,L2R在SAKURA上的表现优于所有对比的LALM,在MMAU和MMAR上则落后6-12个百分点,尽管其训练的参数数量少约1400倍,使用的音频数据也少几个数量级。不过,由于任何LLM都可作为阅读器,我们表明更强的阅读器无需重新训练任何音频组件即可缩小这一差距。添加新领域仅需一个小型头:在每个物种仅5个标注片段的情况下,它在相同片段上的表现比LALM的QLoRA微调高出13-26个百分点。
英文摘要
Large audio-language models (LALMs) fuse an audio encoder into a large language model (LLM) through multi-stage training. This coupling means that a new domain or a stronger LLM requires retraining, and their answers cannot be traced to what the model heard: a chain-of-thought is a post-hoc account. We propose Listen-to-Reason (L2R), an interpretable-by-design pipeline that passes audio to the LLM through an explicit, human-readable tree: small heads on frozen expert encoders map each chunk of a clip to semantically meaningful nodes on the tree (for speech, music and environmental sound), and a frozen text-only LLM answers from these nodes and an ASR transcript without hearing the clip. Every answer can therefore be traced to the nodes and transcript it read, and the nodes are causal: replacing the deciding node with a distractor overturns 78% of correct answers on SAKURA. With a 7B reader, L2R outperforms all LALMs we compare against on SAKURA and trails them by 6-12 points on MMAU and MMAR, despite training about 1,400x fewer parameters on orders of magnitude less audio data. However, because any LLM can serve as the reader, we show that a stronger reader narrows this gap without retraining any audio component. A new domain is added with one small head: with five labelled clips per species, it outperforms QLoRA fine-tuning of an LALM on the same clips by 13-26 points.
发表机构
- Concordia University(康考迪亚大学)
- Mila – Quebec AI Institute(米拉-魁北克人工智能研究所)
- Laval University(拉瓦尔大学)
机构由 AI 辅助整理,请以论文原文为准。