arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向音频语言模型中声学基础的对策略蒸馏奖励倾斜

Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models

Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji

arXiv 2609.28778首次发表:更新:

发表机构

NEC Labs(NEC实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RT-OPD,通过音频输入的对数概率对比重塑教师分布,增强音频语言模型的声学基础,在多个基准上优于基线,3B模型MMAU达72.72%。

AI 中文摘要

音频语言模型(ALMs)可能利用文本捷径来回答问题,而忽略声学证据,从而削弱了音频理解能力。对策略蒸馏(OPD)通过使用教师预测来监督学生生成的响应,从而训练紧凑型ALMs,但并未明确区分声学支持与语言可预测性。我们提出了奖励倾斜对策略蒸馏(RT-OPD)以增强声学基础。给定相同的问题和学生生成的文本,一个冻结的教师模型在有无音频输入的情况下预测下一个词元。它们的对数概率对比定义了一个奖励,该奖励重塑教师分布以进行反向KL蒸馏,强调音频提供的额外证据。在两个紧凑型学生模型和三个基准测试中,RT-OPD始终优于Vanilla OPD。使用静音和替换音频的实验进一步表明,RT-OPD增强了学生对声学证据的依赖。我们的3B模型在MMAU上达到了72.72%的准确率,这是所比较的3B模型中最高的,并且与几个7B和8B模型相当。代码和模型检查点可在以下https URL获取。

英文摘要

Audio-language models (ALMs) can exploit textual shortcuts to answer questions while overlooking acoustic evidence, weakening audio understanding. On-policy distillation (OPD) trains compact ALMs by supervising student-generated responses with teacher predictions, but does not explicitly distinguish acoustic support from linguistic predictability. We propose Reward-Tilted On-Policy Distillation (RT-OPD) to strengthen acoustic grounding. Given the same question and student-generated text, a frozen teacher predicts the next token with and without audio inputs. Their log-probability contrast defines a reward that reshapes the teacher distribution for reverse-KL distillation, emphasizing the additional evidence provided by audio. Across two compact students and three benchmarks, RT-OPD consistently outperforms Vanilla OPD. Experiments with silenced and replacement audio further suggest that RT-OPD strengthens the student's reliance on acoustic evidence. Our 3B model achieves 72.72% accuracy on MMAU, the highest among the compared 3B models and competitive with several 7B and 8B models. Code and model checkpoints are available at https://github.com/KaiyangLi1992/RT-OPD.

Comments5 pages, submitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑