发表机构
Inference Matter Labs(推理物质实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出轻量级多模态模型Qwen-MusicAVQA-7B,通过连接冻结的Whisper编码器与Qwen2-VL-7B-Instruct,在MUSIC-AVQA等基准上实现高准确率,训练成本低,且发现音频时间信息保留程度影响下游问答准确率。
AI 中文摘要
将音频添加到视觉-语言模型的常用方法是训练或适配一个大型全模态系统。我们表明,一种轻量级替代方案对于音乐音频-视觉问答(AVQA)也可以非常有效。Qwen-MusicAVQA-7B通过学习到的线性投影将冻结的Whisper编码器连接到Qwen2-VL-7B-Instruct。同一个冻结编码器通过不同的投影器处理视频的音乐轨道和TTS生成的口语问题,而语言模型通过预训练的自注意力融合视觉帧、音乐和问题音频,无需特定任务的融合网络。在MUSIC-AVQA数据集上,我们的系统在7402个问题的可用视频测试子集上,三个独立训练种子的准确率达到96.0%±3.9%。我们的核心发现是,下游准确率与音频表示保留的细粒度局部时间信息的多少相关。在匹配的32-token比较中,步长池化的Whisper帧序列优于扩展到相同预算的全局池化PANNs表示,优势达26个百分点,尽管PANNs接收的音频至少与Whisper一样多,且使用的投影器大得多。这种效果并非简单的序列与向量差异:仅在Whisper内部,在固定token预算下降低时间分辨率会导致相当的性能损失。在匹配的数据和输入下,微调后的Qwen2.5-Omni-7B达到80.9%的准确率,而我们的30秒变体达到95.9%;由于这些系统在骨干和适配方式上存在差异,这是一个系统级别的比较。在重新表述的MUSIC-AVQA-R基准的采样头和尾分割上,准确率仍然很高,分别为96.5%和95.6%。由于两个编码器都保持冻结且音乐特征被缓存,整个适配的训练成本很低:完整的两阶段AVQA运行在单个A100 80GB上大约需要5小时,本文报告的所有运行都可在该单个GPU上完成。
英文摘要
A common approach to adding audio to a vision-language model is to train or adapt a large omni-modal system. We show that a lightweight alternative can be highly effective for music audio-visual question answering (AVQA). Qwen-MusicAVQA-7B connects a frozen Whisper encoder to Qwen2-VL-7B-Instruct through learned linear projections. The same frozen encoder processes both the video's music track and a TTS-spoken question through separate projectors, while the language model fuses visual frames, music, and question audio through pretrained self-attention, with no task-specific fusion network. On MUSIC-AVQA, our system reaches 96.0% +/- 3.9% accuracy across three independent training seeds on the 7,402-question available-video test subset. Our central finding is that downstream accuracy tracks how much fine-grained local temporal information the audio representation preserves. In a matched 32-token comparison, a stride-pooled Whisper frame sequence outperforms a globally pooled PANNs representation expanded to the same budget by 26 percentage points, even though PANNs sees at least as much audio and uses a far larger projector. The effect is not simply sequence versus vector: within Whisper alone, reducing temporal resolution at a fixed token budget costs a comparable amount. Under matched data and inputs, fine-tuned Qwen2.5-Omni-7B reaches 80.9%, against 95.9% for our 30 s variant; because the systems differ in backbone and adaptation, this is a system-level comparison. Accuracy remains high on sampled head and tail splits of the rephrased MUSIC-AVQA-R benchmark (96.5% and 95.6%). Because both encoders stay frozen and the music features are cached, the entire adaptation is cheap to train: the complete two-stage AVQA run takes approximately 5 hours on a single A100 80GB, and every run reported here fits on that one GPU.
Comments24 pages, 1 figure, 8 tables. Code: https://github.com/MKDehdashti/Qwen2-vl-audio Checkpoints: https://huggingface.co/MayaKD/qwen2-vl-audio