arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23092cs.SD

面向推理的后训练与推理时LoRA重缩放:针对音频相关问答任务

Reasoning-Oriented Post-Training and Inference-Time LoRA Rescaling for Audio-Dependent Question Answering

Weiteng Hu, Yin Cao, Jun Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对音频相关问答任务,提出面向推理的LoRA后训练与推理时重缩放方法,在Qwen和MOSS-Audio模型上验证了有效性,提交系统在挑战赛中获总体第三、轻量级系统第二。

中文摘要 AI 辅助

音频相关问答(ADQA)要求大型音频语言模型(LALMs)回答正确答案依赖给定音频内容的问题。成功的ADQA需要准确的音频感知、问题相关证据的识别以及跨模态推理。利用DCASE 2026任务5的官方ADQA数据集,我们针对Qwen2.5-Omni和MOSS-Audio-8B-Thaming两种模型,研究了面向推理的低秩适配(LoRA)后训练与推理时LoRA重缩放方法。我们引入了结构化思维链(CoT)框架,将推理过程分解为问题分析、问题类型、音频证据和推理四个环节。随后,我们分析了任务特定的LoRA适配对两种主干模型的影响,并进一步探索了训练后LoRA适配器的推理时重缩放。在开发集上的实验显示出明显的主干依赖行为:在我们的监督微调配置下,后训练提升了基于Qwen的系统,但大幅降低了MOSS-Audio的性能;适度的LoRA重缩放进一步将最佳Qwen系统的Top-1准确率从58.93%提升至61.05%,并部分恢复了微调后MOSS-Audio模型的性能,而最佳MOSS-Audio系统达到了67.70%的Top-1准确率。我们提交的系统在挑战赛中总体排名第三,且在参数规模小于10B的轻量级系统中排名第二。

英文摘要

Audio-Dependent Question Answering (ADQA) requires Large Audio-Language Models (LALMs) to answer questions whose correct answers depend on the given audio content. Successful ADQA requires accurate audio perception, identification of question-relevant evidence, and cross-modal reasoning. Using the official ADQA dataset of DCASE 2026 Task 5, we investigate reasoning-oriented post-training with Low-Rank Adaptation (LoRA) and inference-time LoRA rescaling for both Qwen2.5-Omni and MOSS-Audio-8B-Thinking. We introduce a structured Chain-of-Thought (CoT) framework that decomposes the reasoning process into question analysis, question type, audio evidence, and reasoning. We then analyze how task-specific LoRA adaptation affects the two backbones and further explore inference-time rescaling of trained LoRA adapters. Experiments on the development set reveal markedly backbone-dependent behavior: post-training improves the Qwen-based systems but substantially degrades MOSS-Audio under our supervised fine-tuning configuration. Moderate LoRA rescaling further improves the best Qwen system's top-1 accuracy from 58.93% to 61.05% and partially restores the performance of the fine-tuned MOSS-Audio models, while the best MOSS-Audio system achieves 67.70% top-1 accuracy. Our submitted systems ranked third overall and second among lightweight systems under 10B parameters in the challenge.

↑