arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

X³-OPD:通过策略对齐将推理能力提炼到大型音频-语言模型中

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment

Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin

arXiv 2607.21550首次发表:更新:

发表机构

Tencent Hunyuan; Zhejiang University(腾讯混元; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大型音频-语言模型逻辑推理能力不足,提出X³-OPD跨模态策略蒸馏框架,借助教师模型指导与构建的三层对称语料库训练学生模型,实验证明该方法能提升音频推理及思维链质量,还能保留模型域转移下的能力。

AI 中文摘要

虽然大型音频-语言模型在听觉感知方面取得了显著进展,但在深度逻辑推理方面仍落后于基于文本的大型语言模型,主要原因是高质量音频推理数据稀缺。为弥合这一差距,我们提出了X³-OPD,这是一个跨模态策略蒸馏框架,将推理能力从强大的文本教师模型转移到音频-语言学生模型。训练期间,学生模型根据自身声学感知生成推理轨迹,教师模型使用匹配的文本输入和验证答案提供令牌级指导。我们还构建了一个三层对称语料库。实验表明,X³-OPD显著提高了基于音频的推理和思维链质量,同时在很大程度上保留了模型在域转移下的现有能力。

英文摘要

While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training, the student generates reasoning trajectories conditioned on its own acoustic perception, while the teacher provides token-level guidance using matched textual inputs and verified answers. We further construct a three-tier symmetric corpus covering textual reasoning rendered into speech, audio-event reasoning grounded in complex acoustic scenes, and spoken-dialogue reasoning involving paralinguistic cues. This design extends cross-modal distillation beyond textually recoverable content to reasoning grounded in non-linguistic events, prosody, and conversational context. Experiments on MMSU, MMAU, BIG Bench Audio, and MMAR demonstrate that X$^3$-OPD substantially improves audio-grounded reasoning and chain-of-thought quality while largely preserving the model's existing capabilities under domain shift.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑